Audio processing method, apparatus, device, storage medium and program product

CN116034423BActive Publication Date: 2026-08-28GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280004371.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2026-08-28
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

[0004]在上述相关技术中,用户只能采用自己录音得到的音频进行音频制作,制作得到的音频内容较为单一

Benefits of technology

[0016]通过提取第一音频文件的音频特征,并基于第一音频文件的音频特征、和用户的声学模型,将该用户的声学特征与第一音频文件融合,生成具有该用户音色的第二音频文件,实现了对音频进行音色修改的功能,从而提升了音频内容的丰富性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116034423B_ABST
    Figure CN116034423B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio processing method, device, equipment, storage medium and program product, relating to the technical field of audio. The method comprises: obtaining a first audio file (110); extracting an audio feature of the first audio file (120); processing the audio feature through an acoustic model of a first user to generate a second audio file; wherein the acoustic model of the first user is a model learning acoustic features of the first user, and the second audio file has a timbre of the first user (130). The technical solution provided by the embodiments of the present application can improve the richness of audio content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio technology, and in particular to an audio processing method, apparatus, device, storage medium, and program product. Background Technology

[0002] Currently, with the development of audio technology, audio processing methods are becoming increasingly diverse.

[0003] In related technologies, users can record, tune, and play back the audio they produce using an audio production application.

[0004] In the aforementioned technologies, users can only use their own recorded audio to create audio content, resulting in relatively limited audio content. Summary of the Invention

[0005] This application provides an audio processing method, apparatus, device, storage medium, and program product, which can enhance the richness of audio content. The technical solution is as follows:

[0006] According to one aspect of the embodiments of this application, an audio processing method is provided, the method comprising:

[0007] Displays information related to the first audio file;

[0008] In response to a timbre creation instruction for the first audio file, a second audio file generated from the first audio file using the first user's acoustic model is displayed; wherein the first user's acoustic model is a model that has learned the acoustic features of the first user, and the second audio file has the timbre of the first user.

[0009] According to one aspect of the embodiments of this application, an audio processing apparatus is provided, the apparatus comprising:

[0010] The information display module is used to display relevant information about the first audio file;

[0011] The file display module is used to respond to the timbre production instruction for the first audio file and display a second audio file generated from the first audio file by the acoustic model of the first user; wherein the acoustic model of the first user is a model that has learned the acoustic features of the first user, and the second audio file has the timbre of the first user.

[0012] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described audio processing method.

[0013] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the above-described audio processing method.

[0014] According to one aspect of the embodiments of this application, a computer program product is provided, which is loaded and executed by a processor to implement the above-described audio processing method.

[0015] The technical solutions provided in this application embodiment may have the following beneficial effects:

[0016] By extracting the audio features of the first audio file and, based on the audio features of the first audio file and the user's acoustic model, fusing the user's acoustic features with the first audio file, a second audio file with the user's timbre is generated. This achieves the function of modifying the timbre of the audio, thereby enhancing the richness of the audio content.

[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of an audio processing method provided in one embodiment of this application;

[0020] Figure 2 This is a schematic diagram of phonemes provided in one embodiment of this application;

[0021] Figure 3 This is a schematic diagram of an acoustic model provided in one embodiment of this application;

[0022] Figure 4 This is a block diagram of an audio processing apparatus provided in one embodiment of this application;

[0023] Figure 5 This is a block diagram of an audio processing apparatus provided in another embodiment of this application;

[0024] Figure 6 This is a block diagram of a computer device provided in one embodiment of this application. Detailed Implementation

[0025] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of methods consistent with some aspects of this application as detailed in the appended claims.

[0026] The method provided in this application can be executed by a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. This computer device can be a terminal such as a PC (Personal Computer), tablet computer, smartphone, wearable device, or intelligent robot; or it can be a server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0027] The technical solution of this application will be described and illustrated below through several embodiments.

[0028] Please refer to Figure 1 This document illustrates a flowchart of an audio processing method provided in an embodiment of this application. In this embodiment, the method is primarily illustrated by its application in the computer device described above. The method may include the following steps (110-130):

[0029] Step 110: Obtain the first audio file.

[0030] In some embodiments, the first audio file may be an audio file of the type of song, dubbing, poetry recitation, audiobook, radio drama, etc.

[0031] In some embodiments, one or more first audio files are acquired. That is, tone production can be performed on a single audio file; tone production can also be performed on multiple audio files simultaneously, thereby improving tone production efficiency.

[0032] In some embodiments, the first audio file may be an audio file obtained through wired or wireless transmission (e.g., network connection). In some embodiments, the method is applied in a target application of a terminal device (e.g., a client of the target application). The target application may be an audio-type application, such as a music production application, an audio playback application, an audio live streaming application, a karaoke application, etc., which is not specifically limited in the embodiments of the present application. The target application may also be any application having an audio processing function, such as a social application, a payment application, a video application, a shopping application, a news application, a game application, etc. In some embodiments, the first audio file may be an audio file recorded and / or produced by the client of the target application.

[0033] Step 120, extracting audio features of the first audio file.

[0034] In some embodiments, the first audio file includes voice content uttered by any user, and audio features of the voice content uttered by the user are extracted from the first audio file.

[0035] In some embodiments, the audio features include at least one of the following:

[0036] phoneme features, configured to characterize phoneme information of audio content in the first audio file;

[0037] pitch features, configured to characterize pitch information of audio content in the first audio file.

[0038] Wherein, a phoneme refers to the smallest speech unit divided according to the natural properties of speech, and is the smallest linear speech unit divided from the perspective of sound quality. A phoneme is a concrete physical phenomenon. Analyzing based on the articulation movement in a syllable, one movement constitutes one phoneme. In some embodiments, phonemes are divided into two categories: vowels and consonants. For example, the Chinese syllable "ā" has only one phoneme, "ài" has two phonemes, and "dài" has three phonemes. In some embodiments, the phoneme information includes the phonemes contained in the audio content of the first audio file and the pronunciation duration of each phoneme, and these features collectively constitute the phoneme features. For example, some people pronounce more fully, so at a normal speech rate, the pronunciation duration of the vowel corresponding phoneme is relatively long; for another example, some people speak faster with shorter pronunciation, so the duration of each phoneme is relatively short; for another example, affected by physiological factors or living environment, some people have difficulty pronouncing certain phonemes (such as "h", "n", etc.).

[0039] In some embodiments, as Figure 2 shown, each phoneme may be represented by a phoneme block, and the length of the phoneme block is used to represent the pronunciation duration of the corresponding phoneme; for example, the length a1 of phoneme block 21 is used to represent the pronunciation duration of phoneme a.

[0040] Pitch refers to the highness or lowness of a sound, and it is determined by the frequency and wavelength of the sound wave. The higher the frequency and the shorter the wavelength, the higher the pitch; conversely, the lower the frequency and the longer the wavelength, the lower the pitch.

[0041] In some embodiments, the audio features may further include energy features, aerophone features, tension features, etc., of the audio content in the first audio file, and this application does not limit this. Energy features can be used to indicate the volume / loudness of the audio content in the first audio file; aerophone refers to a vocalization method where the vocal cords do not vibrate or vibrate almost not, and aerophone features can indicate the regularity or rhythm of the user's aerophone pronunciation; tension features refer to the variation characteristics between low and high notes, and between soft and loud notes, of the audio content in the first audio file.

[0042] Step 130: Process the audio features using the acoustic model of the first user to generate a second audio file; wherein, the acoustic model of the first user is a model that has learned the acoustic features of the first user, and the second audio file has the timbre of the first user.

[0043] In some embodiments, the acoustic characteristics of the first user include the timbre characteristics of the first user. Timbre refers to the sonic characteristics of different sounds, which are physically manifested in the waveform characteristics of sound waves; therefore, timbre can also be called voiceprint characteristics. The timbre of different people's voices is different.

[0044] In some embodiments, a model learning the acoustic features of a first user is used to process the audio features of a first audio file to generate a second audio file. That is, the timbre of the first user is fused with the audio features (such as phoneme features, pitch features, etc.) of the first audio file to generate a second audio file that combines the timbre of the first user with the phoneme and pitch features of the first audio file.

[0045] In some embodiments, step 130 further includes: processing the audio features using the acoustic model of the first user to generate a mel spectrogram; and generating a second audio file based on the mel spectrogram. Research shows that human perception of sound frequency is not linear, and that low-frequency signals are perceived more readily than high-frequency signals. For example, people can easily perceive the difference between 500 and 1000 Hz (Hertz), but find it difficult to distinguish between 7500 and 8000 Hz. The Mel scale, proposed to address this situation, is a non-linear transformation of sound frequency. For signals (such as sound signals) measured in Mel scale units, it can simulate the linear perception of changes in sound signals by humans.

[0046] In some embodiments, the Mel spectrum can also be replaced with other feasible spectra, and this application does not specifically limit this.

[0047] In some embodiments, such as Figure 3 As shown, the acoustic model 30 includes an encoder 31 and a decoder 32; the audio features are processed using the acoustic model of the first user to generate a Mel spectrum, including the following steps:

[0048] 1. The phoneme features in the audio features are processed by encoder 31 to obtain encoded phoneme features; wherein, the phoneme features are used to characterize the phoneme information of the audio content in the first audio file;

[0049] 2. The encoded phoneme features are fused with the pitch features in the audio features to obtain the fused features;

[0050] 3. The fused features are processed by decoder 32 to obtain the Mel spectrum.

[0051] In some embodiments, encoder 31 acquires phoneme features from audio features, encodes the phoneme features, and obtains encoded phoneme features 33 (also referred to as intermediate layer variables). Optionally, since the pronunciation duration of phonemes is not entirely consistent, during the encoding process, a length adjuster is used to adjust the encoded length of different phoneme features to ensure that the encoded phoneme feature lengths are the same. For example, if the lengths of the phoneme features obtained after preliminary encoding are not uniform, the length of the longest pre-encoded phoneme feature is used as the standard length, and the missing / insufficient parts of other pre-encoded phoneme features relative to the standard length are padded, such as by filling the missing parts with "0", thereby unifying the lengths of all phoneme features and obtaining phoneme features encoded with uniform length. Alternatively, a standard length can be preset, and the missing parts of each phoneme feature relative to the standard length are padded, thereby unifying the lengths of all encoded phoneme features to the standard length. The standard length can be set by those skilled in the art according to actual conditions, and this embodiment does not specifically limit it. Optionally, the standard length is not shorter than the length of the longest phoneme feature after initial encoding.

[0052] In some embodiments, after fusing the encoded phoneme features with the pitch features in the audio features to obtain fused features, the method further includes: extracting slice features of a predetermined length from the fused features; wherein the slice features are used as input to the decoder 32 to obtain the Mel spectrum. That is, the fused features are not used entirely as input to the decoder 32, but rather continuous feature segments of a predetermined length are extracted from them, and these feature segments are sliced ​​to obtain multiple slice features, which are then input to the decoder 32 to obtain the Mel spectrum. In some embodiments, the audio consists of multiple audio frames (i.e., multiple audio segments). Optionally, each audio frame has the same length (i.e., length). The length of one audio frame can be considered as 1, so the length of 100 consecutive audio frames is 100. In some embodiments, each slice feature has the same length (i.e., each slice feature contains the same number of audio frames). For example, the fused feature length is 3000, and multiple continuous slice features are extracted from the fused features and input to the decoder 32, with each slice feature having a length of 500.

[0053] In the above embodiments, only slices of features of a set length are extracted from the fused features for processing, without processing the entire fused features. According to experimental results, this processing has little impact on the model accuracy, thereby saving processing resources and improving the model's processing efficiency while ensuring the accuracy of the acoustic model.

[0054] In some embodiments, the voiceprint features of a first user are obtained; the fused features and the voiceprint features of the first user are processed by a decoder to obtain a Mel spectrum. The audio features of the audio content of the first audio file are then fused with the voiceprint features of the first user to obtain a second audio file that combines the voiceprint features of the first user, the phoneme features, and the pitch features of the first audio file. For singing scenarios, this can result in a song that sounds like it was sung by the first user using the singing style of the singer in the first audio file (i.e., the second audio file), thereby improving the richness of the processed audio file's content.

[0055] In summary, the technical solution provided in this application embodiment integrates the user's acoustic characteristics with the first audio file by using relevant information of the first audio file, timbre production instructions, and the user's acoustic model to generate a second audio file with the user's timbre. This achieves the function of modifying the timbre of the audio, thereby enhancing the richness of the audio content.

[0056] Among some possible implementations, the method also includes:

[0057] 1. Obtain the first user's audio file, which refers to the file obtained by recording the audio content of the first user;

[0058] 2. Using the audio file of the first user, the pre-trained acoustic model is adjusted to obtain the acoustic model of the first user.

[0059] In some embodiments, the first user records an audio file by singing a song, reciting poetry, or dubbing. Based on the first user's audio file, the pre-trained acoustic model is adjusted to obtain the first user's acoustic model.

[0060] In some embodiments, an audio file from a first user is used to adjust a pre-trained acoustic model to obtain an acoustic model for the first user, including:

[0061] (1) Extract the audio features, voiceprint features and standard Mel spectrum corresponding to the first user's audio file;

[0062] (2) Generate a predicted Mel spectrum based on the audio features and voiceprint features corresponding to the first user's audio file using a pre-trained acoustic model;

[0063] (3) Adjust the parameters of the pre-trained acoustic model based on the predicted Mel spectrum and the standard Mel spectrum to obtain the acoustic model of the first user.

[0064] In the above embodiment, the pre-trained acoustic model is fine-tuned using the audio file of the first user. Audio features and voiceprint features extracted from the first user's audio file are input into the pre-trained acoustic model, which outputs the corresponding predicted Mel spectrum. The loss is calculated based on the predicted Mel spectrum and the standard Mel spectrum, and the parameters of the pre-trained acoustic model are adjusted according to the loss calculation results to make its loss function exhibit a gradient descent trend until the fine-tuning of the pre-trained acoustic model is complete, thus obtaining the acoustic model of the first user. This allows for the processing of audio features in the audio file, modifying the voiceprint / timbre of the human speech (such as sung songs, recitations, dubbing, etc.) in the audio file to the voiceprint / timbre of the first user, achieving timbre modification and replacement.

[0065] In some embodiments, the audio features and voiceprint features corresponding to the first user's audio file are preloaded into the GPU (Graphics Processing Unit) memory, thereby eliminating the need to spend more time obtaining the audio features and voiceprint features corresponding to the first user's audio file from elsewhere, thus improving data loading speed and saving model training time.

[0066] In some embodiments, the method further includes: acquiring sample audio files; training an initial acoustic model using the sample audio files to obtain a pre-trained acoustic model. In the above embodiments, audio features, voiceprint features, and standard Mel spectra corresponding to the sample audio files are extracted; the initial acoustic model generates a predicted Mel spectra corresponding to the sample audio files based on the audio features and voiceprint features; then, the parameters of the initial acoustic model are adjusted based on the predicted Mel spectra and the standard Mel spectra corresponding to the sample audio files to obtain the pre-trained acoustic model. The process of training the initial acoustic model using sample audio files to obtain the pre-trained acoustic model can be referred to the relevant content in the above embodiments regarding adjusting the parameters of the pre-trained acoustic model to obtain the acoustic model for the first user, and will not be repeated here.

[0067] The sample audio file can be a relatively large audio file. When the audio file is a song, the sample audio file can include songs sung by celebrities or singers, or songs sung by ordinary people; this application embodiment does not specifically limit this.

[0068] In the above implementation, the pre-trained acoustic model is adjusted based on the first user's audio files to obtain the first user's acoustic model. Since the number of audio files of the first user is small, small sample data can be used to quickly adjust the pre-trained acoustic model, thereby quickly obtaining a personalized acoustic model exclusive to the first user.

[0069] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0070] Please refer to Figure 4 This diagram illustrates a block diagram of an audio processing apparatus according to an embodiment of this application. The apparatus has the functionality to implement the audio processing method example described above; this functionality can be implemented in hardware or by hardware executing corresponding software. The apparatus can be the computer device described above, or it can be mounted on a computer device. The apparatus 400 may include: a file acquisition module 410, a feature extraction module 420, and a file generation module 430.

[0071] The file acquisition module 410 is used to acquire the first audio file.

[0072] The feature extraction module 420 is used to extract the audio features of the first audio file.

[0073] The file generation module 430 is used to process the audio features using the acoustic model of the first user to generate a second audio file; wherein the acoustic model of the first user is a model that has learned the acoustic features of the first user, and the second audio file has the timbre of the first user.

[0074] In some embodiments, the audio features include at least one of the following:

[0075] Phoneme features are used to characterize the phoneme information of the audio content in the first audio file;

[0076] Pitch features are used to characterize the pitch information of the audio content in the first audio file.

[0077] In some embodiments, such as Figure 5 As shown, the file generation module 430 includes a spectrum generation submodule 431 and a file generation submodule 432.

[0078] The spectrum generation submodule 431 is used to process the audio features using the acoustic model of the first user to generate a Mel spectrum.

[0079] The file generation submodule 432 is used to generate the second audio file based on the Mel spectrum.

[0080] In some embodiments, the acoustic model includes an encoder and a decoder; such as Figure 5 As shown, the spectrum generation submodule 431 is used for:

[0081] The encoder processes the phoneme features in the audio features to obtain encoded phoneme features; wherein, the phoneme features are used to characterize the phoneme information of the audio content in the first audio file;

[0082] The encoded phoneme features are fused with the pitch features in the audio features to obtain the fused features;

[0083] The fused features are processed by the decoder to obtain the Mel spectrum.

[0084] In some embodiments, such as Figure 5 As shown, the device 400 further includes a feature extraction module 440.

[0085] The feature extraction module 440 is used to extract slice features of a set length from the fused features; wherein the slice features are used as input to the decoder to obtain the Mel spectrum.

[0086] In some embodiments, such as Figure 5As shown, the device 400 further includes a feature acquisition module 450.

[0087] The feature acquisition module 450 is used to acquire the voiceprint features of the first user.

[0088] The spectrum generation submodule 431 is used to process the fused features and the voiceprint features of the first user through the decoder to obtain the Mel spectrum.

[0089] In some embodiments, such as Figure 5 As shown, the device 400 further includes a model adjustment module 460.

[0090] The file acquisition module 410 is also used to acquire the audio file of the first user, wherein the audio file of the first user refers to the file obtained by recording the audio content of the first user.

[0091] The model adjustment module 460 is used to adjust the pre-trained acoustic model using the audio file of the first user to obtain the acoustic model of the first user.

[0092] In some embodiments, such as Figure 5 As shown, the model adjustment module 460 is used for:

[0093] Extract the audio features, voiceprint features, and standard Mel spectrum corresponding to the first user's audio file;

[0094] The pre-trained acoustic model generates a predicted Mel spectrum based on the audio features and voiceprint features corresponding to the first user's audio file.

[0095] Based on the predicted Mel spectrum and the standard Mel spectrum, the parameters of the pre-trained acoustic model are adjusted to obtain the acoustic model of the first user.

[0096] In some embodiments, the audio features and voiceprint features corresponding to the first user's audio file are preloaded into the GPU memory.

[0097] In some embodiments, such as Figure 5 As shown, the device 400 further includes a model training module 470.

[0098] The file acquisition module 410 is also used to acquire sample audio files.

[0099] The model training module 470 is used to train the initial acoustic model using the sample audio file to obtain the pre-trained acoustic model.

[0100] In summary, the technical solution provided in this application embodiment uses the relevant information of the first audio file, the timbre production instructions, and the user's acoustic model to fuse the user's acoustic characteristics with the first audio file, thereby generating a second audio file with the user's acoustic characteristics, thus enhancing the richness of the audio content.

[0101] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0102] Please refer to Figure 6 This diagram illustrates a structural block diagram of a computer device according to an embodiment of this application. The computer device is used to implement the audio processing method provided in the above embodiments. Specifically:

[0103] The computer device 600 includes a CPU (Central Processing Unit) 601, a system memory 604 including RAM (Random Access Memory) 602 and ROM (Read-Only Memory) 603, and a system bus 605 connecting the system memory 604 and the central processing unit 601. The computer device 600 also includes a basic I / O (Input / Output) system 606 that facilitates information transfer between various components within the computer, and a mass storage device 607 for storing the operating system 613, application programs 614, and other program modules 615.

[0104] The basic input / output system 606 includes a display 608 for displaying information and an input device 609 for user input, such as a mouse or keyboard. Both the display 608 and the input device 609 are connected to the central processing unit 601 via an input / output controller 610 connected to the system bus 605. The basic input / output system 606 may also include the input / output controller 610 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 610 also provides output to a display screen, printer, or other types of output devices.

[0105] The mass storage device 607 is connected to the central processing unit 601 via a mass storage controller (not shown) connected to the system bus 605. The mass storage device 607 and its associated computer-readable media provide non-volatile storage for the computer device 600. That is, the mass storage device 607 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.

[0106] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage media are not limited to the above-mentioned types. The system memory 604 and the mass storage device 607 described above can be collectively referred to as memory.

[0107] According to various embodiments of this application, the computer device 600 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 600 can be connected to a network 612 via a network interface unit 611 connected to the system bus 605, or the network interface unit 611 can be used to connect to other types of networks or remote computer systems (not shown).

[0108] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored therein, which, when executed by a processor, implements the audio processing method described above.

[0109] In an exemplary embodiment, a computer program product is also provided, which is loaded and executed by a processor to implement the above-described audio processing method.

[0110] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0111] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An audio processing method, characterized in that, The method includes: Get the first audio file; Extract audio features from the first audio file. The audio features include phoneme features and pitch features. The phoneme features are used to characterize the phoneme information of the audio content in the first audio file, and the pitch features are used to characterize the pitch information of the audio content in the first audio file. The phoneme features in the audio features are processed by the encoder of the acoustic model of the first user to obtain the encoded phoneme features; wherein, the acoustic model of the first user is a model that has learned the acoustic features of the first user. The encoded phoneme features are fused with the pitch features in the audio features to obtain the fused features; Extract slice features of a set length from the fused features; The slice features are processed by the decoder of the first user's acoustic model to obtain the Mel spectrum; A second audio file is generated based on the Mel spectrum, the second audio file having the timbre of the first user.

2. The method according to claim 1, characterized in that, The method further includes: Obtain the voiceprint features of the first user; The step of processing the fused features through the decoder to obtain the Mel spectrum includes: The decoder processes the fused features and the voiceprint features of the first user to obtain the Mel spectrum.

3. The method according to claim 1, characterized in that, The method further includes: Obtain the audio file of the first user, where the audio file of the first user refers to the file obtained by recording the audio content of the first user; The pre-trained acoustic model is adjusted using the audio file of the first user to obtain the acoustic model of the first user.

4. The method according to claim 3, characterized in that, The step of adjusting the pre-trained acoustic model using the first user's audio file to obtain the first user's acoustic model includes: Extract the audio features, voiceprint features, and standard Mel spectrum corresponding to the first user's audio file; The pre-trained acoustic model generates a predicted Mel spectrum based on the audio features and voiceprint features corresponding to the first user's audio file. Based on the predicted Mel spectrum and the standard Mel spectrum, the parameters of the pre-trained acoustic model are adjusted to obtain the acoustic model of the first user.

5. The method according to claim 4, characterized in that, The audio features and voiceprint features corresponding to the first user's audio file are preloaded into the GPU memory.

6. The method according to claim 5, characterized in that, The method further includes: Obtain sample audio files; The initial acoustic model is trained using the sample audio files to obtain a pre-trained acoustic model.

7. An audio processing device, characterized in that, The device includes: The file acquisition module is used to acquire the first audio file; The feature extraction module is used to extract audio features of the first audio file. The audio features include phoneme features and pitch features. The phoneme features are used to characterize the phoneme information of the audio content in the first audio file, and the pitch features are used to characterize the pitch information of the audio content in the first audio file. The file generation module is used to process the phoneme features in the audio features through the encoder of the acoustic model of the first user to obtain encoded phoneme features; wherein, the acoustic model of the first user is a model that has learned the acoustic features of the first user; the encoded phoneme features are fused with the pitch features in the audio features to obtain fused features; a slice feature of a set length is extracted from the fused features; the slice feature is processed by the decoder of the acoustic model of the first user to obtain a Mel spectrum; and a second audio file is generated based on the Mel spectrum, the second audio file having the timbre of the first user.

8. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, which is loaded and executed by the processor to implement the audio processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the audio processing method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product is loaded and executed by a processor to implement the audio processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for training acoustic conversion model, terminal and storage medium

    CN112992107A

  • Voice conversion method, system, electronic equipment and readable storage medium

    CN113571039A