Audio processing method and apparatus, electronic device, and storage medium
By acquiring the timbre, content, and pitch vectors of audio, and using an audio encoder to generate latent variables and perform mapping and spectrum optimization, the problem of audio enhancement ignoring overall quality in existing technologies is solved, thereby achieving improved audio quality and accurate identification of identity information.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2023-05-31
- Publication Date
- 2026-06-02
AI Technical Summary
In the banking sector, existing technologies for enhancing customer voices primarily focus on intonation while neglecting overall sound quality, resulting in poor audio quality and an inability to accurately determine whether customer identity information is compliant.
By acquiring the timbre, content, and pitch vectors of the audio to be processed, latent variables are generated using an audio encoder, and latent relation mapping and spectrum optimization are performed to improve audio quality.
It improves audio quality, enables more accurate identification of customer information, and enhances the effect of audio enhancement.
Smart Images

Figure CN116543746B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology, and in particular to an audio processing method and apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, in the banking sector, it is necessary to analyze customer voice content to further determine the compliance of customer identity information. One major research area is voice synthesis, which involves the widespread need to enhance customer voices to enable more reliable analysis. In other words, without voice enhancement, it is difficult to perform specific identification and analysis of the voice. The goal of voice enhancement is to improve intonation and tone while maintaining content and timbre. However, most current voice correction solutions typically focus only on intonation, neglecting the overall quality of the voice. This fails to effectively improve audio quality, hindering proper analysis of customer voices and making it difficult to accurately determine the compliance of customer identity information. Therefore, improving the audio quality of voice enhancement has become an urgent technical problem to be solved. Summary of the Invention
[0003] The main objective of this application is to provide an audio processing method, apparatus, electronic device, and storage medium, which aims to improve the audio quality of sound enhancement.
[0004] To achieve the above objectives, a first aspect of this application provides an audio processing method, the method comprising:
[0005] Based on the acquired audio to be processed, determine the first timbre vector, the first content vector, and the first pitch vector corresponding to the audio to be processed;
[0006] Based on a pre-configured audio encoder, a first latent variable corresponding to the audio to be processed is generated according to the first timbre vector, the first content vector, and the first pitch vector;
[0007] Perform latent relation mapping processing on the first latent variable to obtain a second latent variable corresponding to a preset standard audio;
[0008] Align the first content vector with the second pitch vector corresponding to the obtained preset standard audio to obtain the second content vector corresponding to the audio to be processed;
[0009] Based on the audio encoder, according to the first timbre vector, the second content vector, the second pitch vector, and the second latent variable, the initial Mel spectrum of the audio to be processed input into the audio encoder is subjected to spectrum optimization processing to obtain the optimized Mel spectrum of the audio to be processed.
[0010] In some embodiments, determining the first timbre vector, first content vector, and first pitch vector corresponding to the acquired audio to be processed includes:
[0011] From the acquired audio to be processed, the initial Mel spectrum of the audio to be processed and the audio vector corresponding to the audio to be processed are extracted;
[0012] Based on the initial Mel spectrum and the audio vector, the first timbre vector, the first content vector, and the first pitch vector corresponding to the audio to be processed are determined.
[0013] In some embodiments, determining the first timbre vector, first content vector, and first pitch vector corresponding to the audio to be processed based on the initial Mel spectrum and the audio vector includes:
[0014] The initial Mel spectrum is input into the pre-configured timbre encoder and content encoder respectively to obtain the first timbre vector output by the timbre encoder and the first content vector output by the content encoder.
[0015] The audio vector is input into a pre-configured pitch encoder to obtain a first pitch vector output by the pitch encoder.
[0016] In some embodiments, aligning the first content vector with the second pitch vector corresponding to the obtained preset standard audio to obtain the second content vector corresponding to the audio to be processed includes:
[0017] The first content vector is aligned with the second pitch vector corresponding to the preset standard audio by using a dynamic time warping algorithm to obtain the second content vector corresponding to the audio to be processed.
[0018] In some embodiments, the step of performing latent relation mapping processing on the first latent variable to obtain a second latent variable corresponding to a preset standard audio includes:
[0019] The first latent variable is mapped to a third latent variable corresponding to a preset standard audio using a latent relation mapping engine algorithm.
[0020] The third latent variable is subjected to data optimization processing to obtain a second latent variable corresponding to a preset standard audio.
[0021] In some embodiments, the step of performing data optimization processing on the third latent variable to obtain a second latent variable corresponding to a preset standard audio includes:
[0022] The third latent variable is trained using a pre-trained log-likelihood model for maximum likelihood estimation.
[0023] In one embodiment, before generating the first latent variable corresponding to the audio to be processed based on the first timbre vector, the first content vector, and the first pitch vector using the pre-configured audio encoder, the method further includes:
[0024] The pre-configured audio encoder is trained by maximizing the lower bound of evidence and by adversarial learning.
[0025] To achieve the above objectives, a second aspect of this application provides an audio processing apparatus, the apparatus comprising:
[0026] The vector output module is used to determine the first timbre vector, the first content vector, and the first pitch vector corresponding to the acquired audio to be processed.
[0027] The first processing module is used to generate a first latent variable corresponding to the audio to be processed based on a pre-configured audio encoder, according to the first timbre vector, the first content vector, and the first pitch vector.
[0028] The second processing module is used to perform latent relation mapping processing on the first latent variable to obtain a second latent variable corresponding to the preset standard audio.
[0029] The alignment processing module is used to align the first content vector with the second pitch vector corresponding to the obtained preset standard audio to obtain the second content vector corresponding to the audio to be processed.
[0030] An optimization processing module is used to perform spectrum optimization processing on the initial Mel spectrum of the audio to be processed input into the audio encoder based on the audio encoder, according to the first timbre vector, the second content vector, the second pitch vector, and the second latent variable, to obtain the optimized Mel spectrum of the audio to be processed.
[0031] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0032] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0033] The audio processing method, apparatus, electronic device, and storage medium proposed in this application determine the audio parameters corresponding to the audio to be processed, including a first timbre vector, a first content vector, and a first pitch vector, and generate corresponding first latent variables by processing each audio parameter based on an audio encoder. The first latent variables are then converted into latent variables corresponding to professional sound quality to improve the sound quality of the audio to be processed. Furthermore, the first content vector is aligned with the second pitch vector of the professional sound quality to improve the pitch of the audio to be processed. Based on the improved audio parameters, the Mel spectrum of the audio to be processed is optimized to obtain a more aesthetically pleasing Mel spectrum, which is beneficial to improving the audio quality of sound enhancement. Attached Figure Description
[0034] Figure 1 This is a flowchart of an audio processing method provided in one embodiment of this application;
[0035] Figure 2 yes Figure 1 The flowchart of step S101 in the text;
[0036] Figure 3 yes Figure 2 The flowchart of step S202 in the document;
[0037] Figure 4 yes Figure 1 The flowchart of step S102 in the document;
[0038] Figure 5 yes Figure 1 The flowchart of step S103 in the process;
[0039] Figure 6 yes Figure 5 The flowchart of step S502 in the document;
[0040] Figure 7 yes Figure 1 The flowchart of step S104 in the process;
[0041] Figure 8 This is a schematic diagram of the execution flow of an audio processing method provided in one embodiment of this application;
[0042] Figure 9 This is a schematic diagram of the structure of an audio processing device provided in one embodiment of this application;
[0043] Figure 10 This is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0045] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0047] First, let's analyze some of the terms used in this application:
[0048] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0049] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.
[0050] Currently, in banking settings, enhancements can be performed by professional audio engineers with sufficient domain knowledge, using commercial vocal correction tools such as Melodyne 3 and Autotune 4. While most automatic pitch correction tools have proven to be attractive, they may exhibit weak alignment or pitch accuracy, potentially resulting in homogeneous voice styles between tuned and reference recordings. Furthermore, because they typically focus solely on intonation, they easily overlook overall quality, i.e., audio quality and timbre.
[0051] Based on this, embodiments of this application provide an audio processing method, apparatus, electronic device, and storage medium, aiming to improve the audio quality of sound enhancement and enhance sound quality. By introducing an audio encoder as an enhancement system, it is possible to convert intonation and tone while maintaining content and timbre. This is somewhat different from the conversion tasks in related technologies. In fact, conversion belongs to a sub-task of speech conversion. That is to say, the audio processing method provided by embodiments of this application can also be applied to relevant application conditions of speech conversion, but is not limited to them, and has broad application prospects.
[0052] The audio processing method, apparatus, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the audio processing method in the embodiments of this application is described.
[0053] The embodiments of this application can acquire and process relevant data based on financial technology. Here, financial (Artificial Intelligence, AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0054] Financial infrastructure technologies generally include technologies such as sensors, dedicated financial chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. Financial software technologies mainly encompass several areas, including computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0055] The audio processing method provided in this application relates to the field of financial technology. The audio processing method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet computer, laptop computer, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and financial platforms; the software can be an application implementing the audio processing method, but is not limited to the above forms.
[0056] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via communication networks. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0057] It should be noted that in various specific embodiments of this application, when processing is required based on user information, user behavior data, user historical data, and user location information, which are related to user identity or characteristics—for example, when obtaining user-related audio data to be processed—user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards of the relevant countries and regions. In addition, when embodiments of this application require obtaining sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the normal operation of the embodiments of this application obtained.
[0058] Figure 1 This is an optional flowchart of the audio processing method provided in the embodiments of this application. Figure 1The method may include, but is not limited to, steps S101 to S105.
[0059] Step S101: Based on the acquired audio to be processed, determine the first timbre vector, the first content vector, and the first pitch vector corresponding to the audio to be processed;
[0060] Step S102: Based on the pre-configured audio encoder, generate a first latent variable corresponding to the audio to be processed according to the first timbre vector, the first content vector, and the first pitch vector;
[0061] Step S103: Perform latent relation mapping processing on the first latent variable to obtain the second latent variable corresponding to the preset standard audio.
[0062] Step S104: Align the first content vector with the second pitch vector corresponding to the obtained preset standard audio to obtain the second content vector corresponding to the audio to be processed.
[0063] Step S105: Based on the audio encoder, according to the first timbre vector, the second content vector, the second pitch vector, and the second latent variable, the initial Mel spectrum of the audio to be processed input into the audio encoder is subjected to spectrum optimization processing to obtain the optimized Mel spectrum of the audio to be processed.
[0064] Steps S101 to S105 of this application embodiment involve determining the audio parameters corresponding to the audio to be processed, including a first timbre vector, a first content vector, and a first pitch vector, and generating corresponding first latent variables by processing each audio parameter based on the audio encoder. The first latent variables are then converted into latent variables corresponding to professional sound quality to improve the sound quality of the audio to be processed. Furthermore, the first content vector is aligned with the second pitch vector of the professional sound quality to improve the pitch of the audio to be processed. Based on the improved audio parameters, the Mel spectrum of the audio to be processed is optimized to obtain a more aesthetically pleasing Mel spectrum, which is beneficial for improving the audio quality of sound enhancement.
[0065] In some embodiments, steps S101 to S105 divide the audio enhancement task into two parts: pitch correction and pitch improvement. The first part corrects the intonation by aligning the pitch curve of the audio to be processed with the pitch curve of a preset standard audio template, and then combining the aligned curves to resynthesize a new sound sample. The second part improves sound quality by converting the latent variables of amateur sound quality in the audio to be processed into latent variables of professional sound quality through latent relation mapping, thereby improving sound quality. Taking banking services as an example, pitch correction and pitch improvement can enhance a customer's speech to standard speech. This standard speech is then compared with the identity information in the big data system. If a match is found, the customer is a registered customer of the bank, and the bank will process the service. Otherwise, the customer is determined not to be a registered customer and needs to register before being served.
[0066] In step S101 of some embodiments, the acquired audio to be processed can be acquired in real time or in advance, and there is no limitation here; the audio to be processed can be the voice audio of a singer or an ordinary speaker, including a speaker, a speaker in special circumstances, etc., that is, the source of the audio to be processed is not limited.
[0067] In step S101 of some embodiments, timbre, content, and pitch are audio parameters well known in the field of sound. Those skilled in the art can clearly distinguish their meanings and differences, and for the sake of avoiding redundancy, they will not be described in detail here.
[0068] Please see Figure 2 In some embodiments, step S101 may include, but is not limited to, steps S201 to S202:
[0069] Step S201: Extract the initial Mel spectrum of the audio to be processed and the audio vector corresponding to the audio to be processed from the acquired audio to be processed.
[0070] Step S202: Based on the initial Mel spectrum and audio vector, determine the first timbre vector, first content vector, and first pitch vector corresponding to the audio to be processed.
[0071] In this step, by extracting the initial Mel spectrum of the audio to be processed and the corresponding audio vector, the sound state of the audio to be processed can be known, so as to further determine the first timbre vector, the first content vector, and the first pitch vector of the audio to be processed based on the initial Mel spectrum and the audio vector.
[0072] In step S201 of some embodiments, the initial Mel spectrum is the actual Mel spectrum corresponding to the audio to be processed. The initial Mel spectrum is unprocessed and unoptimized, and belongs to the characteristic Mel spectrum corresponding to the audio to be processed. It should be noted that the audio vector may, but is not limited to, correspond to the pitch feature of the audio to be processed.
[0073] Please see Figure 3 In some embodiments, step S202 may include, but is not limited to, steps S301 to S302:
[0074] Step S301: Input the initial Mel spectrum into the pre-configured timbre encoder and content encoder respectively to obtain the first timbre vector output by the timbre encoder and the first content vector output by the content encoder.
[0075] Step S302: Input the audio vector into the pre-configured pitch encoder to obtain the first pitch vector output by the pitch encoder.
[0076] In this step, different encoders are set to ensure that the corresponding first timbre vector, first content vector, and first pitch vector can be output separately, so that the output of each vector will not have a significant impact on each other. This helps to determine the first timbre vector, first content vector, and first pitch vector corresponding to the audio to be processed more reliably as a whole.
[0077] In steps S301 and S302 of some embodiments, the types and structures of the timbre encoder, content encoder and pitch encoder can be set according to specific application scenarios, and are not limited here. Specific embodiments will be given below for description, and will not be repeated here to avoid redundancy.
[0078] In step S102 of some embodiments, the audio encoder can be selected and set according to the specific application scenario. For example, it can be set as a Conditional Variational Autoencoder (CVAE), but is not limited to this. The Conditional Variational Autoencoder specifically includes a Variational Auto Encoder (VAE) encoder and a VAE decoder, which are used to realize the complete processing of the Mel spectrum.
[0079] Please see Figure 4 In some embodiments, step S401 may be included, but is not limited to, before step S102:
[0080] Step S401: Perform maximum evidence lower bound training and adversarial learning training on the pre-configured audio encoder.
[0081] In this step, by maximizing the lower bound of evidence and conducting adversarial learning training on the pre-configured audio encoder, the encoding performance of the audio encoder can be optimized, and the robustness of the audio encoder can be improved.
[0082] In step S401 of some embodiments, the specific application forms of maximizing the lower bound of evidence training and adversarial learning training can be various, such as using a neural network, etc., which are not limited here, and this part is well known to those skilled in the art, so it will not be described in detail.
[0083] Please see Figure 5 In some embodiments, step S103 may include, but is not limited to, steps S501 to S502:
[0084] Step S501: Using the latent relation mapping engine algorithm, the first latent variable is mapped to a third latent variable corresponding to the preset standard audio.
[0085] Step S502: Perform data optimization processing on the third latent variable to obtain the second latent variable corresponding to the preset standard audio.
[0086] In this step, the second and third latent variables corresponding to the preset standard audio are the latent variables of professional sound quality. In other words, by using the latent relation mapping engine algorithm to map the first latent variable to the third latent variable corresponding to the preset standard audio, the technical effect of improving sound quality can be achieved. Furthermore, by performing data optimization processing on the third latent variable to obtain the second latent variable corresponding to the preset standard audio, the latent variable can be further optimized to make its application performance better.
[0087] In step S501 of some embodiments, the latent relation mapping engine algorithm is used to change the form of variables, that is, to modify ordinary variables into latent variables. The specific form of the latent relation mapping engine algorithm can be various. Those skilled in the art can select the appropriate latent relation mapping engine algorithm for application according to the specific application scenario. There is no limitation here.
[0088] Please see Figure 6 In some embodiments, step S502 may include, but is not limited to, step S601:
[0089] Step S601: The third latent variable is trained using a pre-trained log-likelihood model for maximum likelihood estimation.
[0090] In this step, the obtained third latent variable is trained by maximum likelihood estimation using a pre-trained log-likelihood model. This allows for further optimization of the third latent variable. In other words, the probability model corresponding to the maximum likelihood can be used to find a phylogenetic tree that can generate the observed data with a high probability. It can be understood that the maximum likelihood method is a representative of a class of phylogenetic tree reconstruction methods that are entirely based on statistics, but it is not limited to this. Similar training methods can also be used to optimize the third latent variable. This is not limited here.
[0091] Please see Figure 7 In some embodiments, step S104 includes, but is not limited to, step S701:
[0092] Step S701: The first content vector is aligned with the second pitch vector corresponding to the preset standard audio using a dynamic time warping algorithm to obtain the second content vector corresponding to the audio to be processed.
[0093] In this step, the alignment of the first content vector and the second pitch vector is performed using the Dynamic Time Warping (DTW) algorithm, which can more accurately align the content vector with the professional pitch vector, thereby improving the robustness of the alignment methods in related technologies.
[0094] In step S701 of some embodiments, the time series data can contain various similarity or distance functions, one of which is the DTW algorithm. This algorithm, based on dynamic programming, solves the template matching problem for speech patterns with varying lengths and is used for isolated word recognition. While the HMM algorithm requires a large amount of speech data during training and involves repeated calculations to obtain model parameters, the DTW algorithm requires almost no additional computation during training, thus significantly reducing training overhead. Therefore, the DTW algorithm can be widely used in isolated word speech recognition.
[0095] To better illustrate the working principle and content of the above embodiments, a specific example is given below.
[0096] Example 1:
[0097] Please see Figure 8 The diagram illustrates the execution flow of an audio processing method according to an embodiment of this application.
[0098] like Figure 8 As shown, the execution of this audio processing method mainly consists of two stages, as detailed below:
[0099] The first stage mainly consists of a pitch encoder, a content encoder, and a timbre encoder. The pitch encoder can, but is not limited to, consist of three convolutional layers, and it receives external audio vectors (i.e.,...). Figure 8 The vector "Pitch" shown in the first stage or the "Pitch" shown in the second stage a "Pitch" p The content encoder and timbre encoder can be designed according to the following requirements: For example, given a recording of singing, in order to obtain its content vector, a Conform-based Automatic Speech Recognition (ASR) model can be trained using the speech and singing data, and the hidden state as language content information, also known as the speech post-column, can be extracted from the output of the ASR model (considered as the content encoder). For obtaining vocal timbre, the open-source API similarity encoder (similarity blyzer8) is used as the timbre encoder. This is a deep learning model designed for speaker verification that can extract the singer's identity information. Under the execution of the above process, based on the pitch, content, and timbre conditions extracted from the input by the pitch encoder, content encoder, and timbre encoder, the Mel spectrogram of the input (i.e., the "initial Mel spectrum") is reconstructed through the CVAE backbone, and the CVAE is optimized by maximizing the lower bound of evidence and adversarial learning.
[0100] In the second stage, the latent variable z is first inferred based on amateur conditions (i.e., corresponding to the "audio to be processed"). a Its basic approach is the same as the first stage, except that it adds a latent variable z based on the audio parameters generated by each encoder. a The steps are straightforward and will not be elaborated here; secondly, the amateur content vector z is processed using the DTW algorithm. a (Right now Figure 8 The second phase shows the "Pitch" a ") and professional pitch content vector z p (Right now Figure 8 The second phase shows the "Pitch" p Align z using a latent mapping algorithm. a Mapping to z p Finally, the professional pitch vector, the aligned amateur content vector, and the amateur timbre vector are mixed to obtain a new condition that corresponds to the z-axis of the VAE encoder mapped by the VAE decoder. p Together, they can generate a new, aesthetically pleasing Mel spectrum.
[0101] As can be seen, aligning amateur content vectors with professional pitch vectors using the DTW algorithm can improve the robustness of existing time warp methods. The audio processing method of this application can not only correct the pitch of amateur recordings, but also generate audio with high audio quality and improved timbre. In this process, CVAE is used as the backbone for generating high-quality audio, and the latent representation of timbre is learned, resulting in better processing effect.
[0102] To correct intonation, aligning amateur recordings with professional pitch curves and then combining them to resynthesize a new singing sample can significantly reduce errors in the audio processing process and expand the application scenarios of audio processing. Furthermore, based on the latent space-based latent mapping algorithm, the latent variables of amateur audio quality can be converted into the latent variables of professional audio quality, thereby achieving the technical effect of improving audio quality.
[0103] Please see Figure 9 This application also provides an audio processing apparatus that can implement the above-described audio processing method. The apparatus includes:
[0104] The vector output module is used to determine the first timbre vector, the first content vector, and the first pitch vector corresponding to the acquired audio to be processed.
[0105] The first processing module is used to generate a first latent variable corresponding to the audio to be processed based on a pre-configured audio encoder, according to a first timbre vector, a first content vector, and a first pitch vector.
[0106] The second processing module is used to perform latent relation mapping processing on the first latent variable to obtain the second latent variable corresponding to the preset standard audio.
[0107] The alignment processing module is used to align the first content vector with the second pitch vector corresponding to the obtained preset standard audio to obtain the second content vector corresponding to the audio to be processed.
[0108] The optimization processing module is used to perform spectrum optimization processing on the initial Mel spectrum of the audio to be processed input into the audio encoder based on the first timbre vector, the second content vector, the second pitch vector, and the second latent variable, so as to obtain the optimized Mel spectrum of the audio to be processed.
[0109] The specific implementation of this audio processing device is basically the same as the specific embodiment of the audio processing method described above, and belongs to the same inventive concept, so it will not be described again here.
[0110] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described audio processing method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0111] Please see Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0112] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0113] The memory 902 can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the audio processing method of the embodiments of this application.
[0114] The input / output interface 903 is used to implement information input and output;
[0115] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0116] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);
[0117] The processor 901, memory 902, input / output interface 903, and communication interface 904 communicate with each other within the device via bus 905.
[0118] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described audio processing method.
[0119] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state memory device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0120] The audio processing method, apparatus, electronic device, and storage medium provided in this application determine the audio parameters corresponding to the audio to be processed, including a first timbre vector, a first content vector, and a first pitch vector, and generate corresponding first latent variables by processing each audio parameter based on an audio encoder. The first latent variables are then converted into latent variables corresponding to professional sound quality to improve the sound quality of the audio to be processed. Furthermore, the first content vector is aligned with the second pitch vector of the professional sound quality to improve the pitch of the audio to be processed. Based on the improved audio parameters, the Mel spectrum of the audio to be processed is optimized to obtain a more aesthetically pleasing Mel spectrum, which is beneficial to improving the audio quality of sound enhancement.
[0121] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0122] The foregoing has described specific embodiments of this application; other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily have to follow the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0123] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and computer-readable storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0124] The apparatus, device, computer-readable storage medium and method provided in the embodiments of this application are corresponding. Therefore, the apparatus, device and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device and computer storage medium will not be described again here.
[0125] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement to the methodology cannot be implemented using hardware physical modules.
[0126] For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system onto a PLD themselves, eliminating the need for chip manufacturers to design and fabricate dedicated integrated circuit chips. Furthermore, instead of manually fabricating integrated circuit chips, this programming is now mostly implemented using "logic compiler" software, similar to the software compiler used in program development. The source code before compilation must be written in a specific programming language called a Hardware Description Language (HDL). There is not just one type of HDL, but many, such as:
[0127] ABEL (Advanced Boolean Expression Language); AHDL (Altera Hardware Description Language); Confluence; CUPL (Cornell University Programming Language); HDCal; and JHDL (Java Hardware Description Language); Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, among the technologies in this field, VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog are more commonly used. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using the aforementioned hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0128] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) that can be executed by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers:
[0129] The memory controller, including the ARC 625D, Atmel AT91SAM, Microchip IP address PIC18F26K20, and Silicon Labs C8051F320, can also be implemented as part of the memory's control logic. Those skilled in the art will also recognize that, in addition to implementing the controller as purely computer-readable program code, the same functionality can be achieved by logically programming the method steps, making the controller function as logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers (PLCs), and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the devices included within it for implementing various functions can also be considered structures within that hardware component. Alternatively, the devices for implementing various functions can be considered as both software modules implementing the method and structures within a hardware component.
[0130] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0131] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0132] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0133] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0134] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0135] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0136] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0137] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0138] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0139] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0140] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, A and B simultaneously, or B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.
[0141] The embodiments of this application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. The embodiments of this application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.
[0142] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0143] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. An audio processing method, characterized in that, include: Based on the acquired audio to be processed, determine the first timbre vector, the first content vector, and the first pitch vector corresponding to the audio to be processed; Based on a pre-configured audio encoder, a first latent variable corresponding to the audio to be processed is generated according to the first timbre vector, the first content vector, and the first pitch vector; Perform latent relation mapping processing on the first latent variable to obtain a second latent variable corresponding to a preset standard audio; Based on isolated word recognition, the first content vector is aligned with the second pitch vector corresponding to the preset standard audio to obtain the second content vector corresponding to the audio to be processed. Based on the audio encoder, according to the first timbre vector, the second content vector, the second pitch vector, and the second latent variable, the initial Mel spectrum of the audio to be processed input into the audio encoder is subjected to spectrum optimization processing to obtain the optimized Mel spectrum of the audio to be processed. The step of performing latent relation mapping processing on the first latent variable to obtain a second latent variable corresponding to a preset standard audio includes: The first latent variable is mapped to a third latent variable corresponding to a preset standard audio using a latent relation mapping engine algorithm. The third latent variable is subjected to data optimization processing to obtain a second latent variable corresponding to a preset standard audio.
2. The audio processing method according to claim 1, characterized in that, The step of determining the first timbre vector, first content vector, and first pitch vector corresponding to the acquired audio to be processed includes: From the acquired audio to be processed, the initial Mel spectrum of the audio to be processed and the audio vector corresponding to the audio to be processed are extracted; Based on the initial Mel spectrum and the audio vector, the first timbre vector, the first content vector, and the first pitch vector corresponding to the audio to be processed are determined.
3. The audio processing method according to claim 2, characterized in that, The step of determining the first timbre vector, first content vector, and first pitch vector corresponding to the audio to be processed based on the initial Mel spectrum and the audio vector includes: The initial Mel spectrum is input into the pre-configured timbre encoder and content encoder respectively to obtain the first timbre vector output by the timbre encoder and the first content vector output by the content encoder. The audio vector is input into a pre-configured pitch encoder to obtain a first pitch vector output by the pitch encoder.
4. The audio processing method according to claim 1, characterized in that, The step of aligning the first content vector with the second pitch vector corresponding to the preset standard audio obtained based on isolated word recognition to obtain the second content vector corresponding to the audio to be processed includes: The first content vector is aligned with the second pitch vector corresponding to the preset standard audio by using a dynamic time warping algorithm to obtain the second content vector corresponding to the audio to be processed.
5. The audio processing method according to claim 1, characterized in that, The step of performing data optimization processing on the third latent variable to obtain the second latent variable corresponding to the preset standard audio includes: The third latent variable is trained using a pre-trained log-likelihood model for maximum likelihood estimation.
6. The audio processing method according to claim 1, characterized in that, Before generating the first latent variable corresponding to the audio to be processed based on the first timbre vector, the first content vector, and the first pitch vector, the pre-configured audio encoder further includes: The pre-configured audio encoder is trained by maximizing the lower bound of evidence and by adversarial learning.
7. An audio processing device, characterized in that, The device includes: The vector output module is used to determine the first timbre vector, the first content vector, and the first pitch vector corresponding to the acquired audio to be processed. The first processing module is used to generate a first latent variable corresponding to the audio to be processed based on a pre-configured audio encoder, according to the first timbre vector, the first content vector, and the first pitch vector. The second processing module is used to perform latent relation mapping processing on the first latent variable to obtain a second latent variable corresponding to the preset standard audio. The alignment processing module is used to align the first content vector with the second pitch vector corresponding to the preset standard audio obtained based on isolated word recognition, so as to obtain the second content vector corresponding to the audio to be processed. An optimization processing module is used to perform spectrum optimization processing on the initial Mel spectrum of the audio to be processed input into the audio encoder based on the audio encoder, according to the first timbre vector, the second content vector, the second pitch vector, and the second latent variable, to obtain the optimized Mel spectrum of the audio to be processed. The step of performing latent relation mapping processing on the first latent variable to obtain a second latent variable corresponding to a preset standard audio includes: The first latent variable is mapped to a third latent variable corresponding to a preset standard audio using a latent relation mapping engine algorithm. The third latent variable is subjected to data optimization processing to obtain a second latent variable corresponding to a preset standard audio.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the audio processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the audio processing method according to any one of claims 1 to 6.