Method and apparatus for audio processing
By automating the processing of the original song's audio signal and text information, and combining it with a music style transfer model, automatic music style conversion is achieved, solving the problems of high labor costs and low efficiency in existing technologies, and generating target song audio that is consistent with the style of the reference song.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-29
Smart Images

Figure CN122116854A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular to an audio processing method and apparatus. Background Technology
[0002] As audio technology continues to improve, users are increasingly demanding a wider variety of music styles. For example, a user might want to hear a rock version of a song that was originally folk; this change in music style is called style migration.
[0003] Currently, in related technologies, solutions for music style transfer typically involve professionals using specialized music editing software to edit the song's audio to achieve the desired music style transfer.
[0004] However, manually transferring musical styles requires significant human resources and is inefficient. Summary of the Invention
[0005] This application provides an audio processing method and apparatus that can automatically achieve music style transfer without manual intervention, resulting in higher efficiency. The technical solution is as follows: Firstly, an audio processing method is provided, the method comprising: Extract the beat audio signal, chord audio signal, and vocal audio signal from the first song's audio. The beat audio signal, chord audio signal and vocal audio signal are mixed to obtain the control audio signal of the first song, wherein the control audio signal is used to characterize the chords and beat of the first song; Obtain text information, wherein the text information includes music style indication information; Obtain the musical style features of the second song, wherein the musical style features are used to characterize the musical style of the second song; Based on the control audio signal of the first song, the text information, and the musical style features of the second song, a target song audio is generated. The musical style of the target song audio matches the musical style of the second song and the musical style indicated by the musical style indication information. The chords and rhythms of the target song audio match the chords and rhythms of the first song.
[0006] In one possible implementation, the beat audio signal, chord audio signal, and vocal audio signal are all in the format of waveform audio files (wav).
[0007] In one possible implementation, the extraction of the beat audio signal, chord audio signal, and vocal audio signal of the first song audio includes: Extract the beat audio signal, chord audio signal, and vocal audio signal of the first paragraph of the first song; The control audio signal for the first song obtained by mixing the beat audio signal, chord audio signal, and vocal audio signal includes: The beat audio signal, chord audio signal and vocal audio signal are mixed to obtain the control audio signal corresponding to the first paragraph of the first song. The control audio signal corresponding to the first paragraph is used to represent the chords and beat of the first paragraph, and also to represent the paragraph information of the first paragraph.
[0008] In one possible implementation, obtaining the musical style features of the second song includes: Obtain the musical style feature vector of the entire second song and the audio of the second section; The process of generating the target song audio based on the control audio signal of the first song, the text information, and the musical style features of the second song includes: The control audio signal corresponding to the first paragraph, the text information, and the song audio of the second paragraph are processed by word segmentation to obtain the first word unit token sequence; Insert a specified candidate token at the end of the first token sequence to obtain the second token sequence; The second token sequence is processed into a feature vector to obtain the first feature vector; Replace the feature sequence corresponding to the specified candidate token in the first feature vector with the music style feature vector of the entire second song to obtain the second feature vector; The second feature vector is input into the trained music style transfer model to obtain the target song audio of the first segment. The music style of the target song audio of the first segment matches the music style of the second segment of the second song and the music style indicated by the music style indication information. The chords and rhythms of the target song audio of the first segment match the chords and rhythms of the first segment of the first song.
[0009] In one possible implementation, the text information also includes the lyrics text of the first paragraph of the first song.
[0010] In one possible implementation, the step of generating the target song audio based on the control audio signal of the first song, the text information, and the musical style features of the second song includes: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the audio of the target song.
[0011] In one possible implementation, the step of inputting the control audio signal of the first song, the text information, and the musical style features of the second song into a trained musical style transfer model to obtain the target song audio includes: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the target accompaniment audio. The target accompaniment audio and the vocal audio of the first song are mixed to obtain the target song audio.
[0012] In one possible implementation, the step of inputting the control audio signal of the first song, the text information, and the musical style features of the second song into a trained musical style transfer model to obtain the target accompaniment audio includes: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the third token sequence. The third token sequence is scored. If the score of the third token sequence is greater than a threshold, the target accompaniment audio is generated based on the third token sequence.
[0013] Secondly, an audio processing apparatus is provided, the apparatus comprising: The acquisition module is used to extract the beat audio signal, chord audio signal, and vocal audio signal of the first song; mix the beat audio signal, chord audio signal, and vocal audio signal to obtain the control audio signal of the first song, wherein the control audio signal is used to characterize the chords and beat of the first song; acquire text information, wherein the text information includes music style indication information; and acquire the music style features of the second song, wherein the music style features are used to characterize the music style of the second song. The inference module is used to generate song audio based on the control audio signal of the first song, the text information, and the musical style features of the second song to obtain target song audio. The musical style of the target song audio matches the musical style of the second song and the musical style indicated by the musical style indication information. The chords and rhythms of the target song audio match the chords and rhythms of the first song.
[0014] In one possible implementation, the beat audio signal, chord audio signal, and vocal audio signal are all in the format of waveform audio files (wav).
[0015] In one possible implementation, the acquisition module is configured to: Extract the beat audio signal, chord audio signal, and vocal audio signal of the first paragraph of the first song; The beat audio signal, chord audio signal and vocal audio signal are mixed to obtain the control audio signal corresponding to the first paragraph of the first song. The control audio signal corresponding to the first paragraph is used to represent the chords and beat of the first paragraph, and also to represent the paragraph information of the first paragraph.
[0016] In one possible implementation, the acquisition module is configured to: Obtain the musical style feature vector of the entire second song and the audio of the second section; The inference module is used for: The control audio signal corresponding to the first paragraph, the text information, and the song audio of the second paragraph are processed by word segmentation to obtain the first word unit token sequence; Insert a specified candidate token at the end of the first token sequence to obtain the second token sequence; The second token sequence is processed into a feature vector to obtain the first feature vector; Replace the feature sequence corresponding to the candidate placeholder in the first feature vector with the music style feature vector of the entire second song to obtain the second feature vector; The second feature vector is input into the trained music style transfer model to obtain the target song audio of the first segment. The music style of the target song audio of the first segment matches the music style of the second segment of the second song and the music style indicated by the music style indication information. The chords and rhythms of the target song audio of the first segment match the chords and rhythms of the first segment of the first song.
[0017] In one possible implementation, the text information also includes the lyrics text of the first paragraph of the first song.
[0018] In one possible implementation, the acquisition module is configured to: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the audio of the target song.
[0019] In one possible implementation, the inference module is used for: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the target accompaniment audio. The target accompaniment audio and the vocal audio of the first song are mixed to obtain the target song audio.
[0020] In one possible implementation, the inference module is used for: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the third token sequence. The third token sequence is scored. If the score of the third token sequence is greater than a threshold, the target accompaniment audio is generated based on the third token sequence.
[0021] Thirdly, a computing device is provided, characterized in that the computing device includes a processor and a memory, the memory storing at least one instruction, the instruction being loaded and executed by the processor to perform the operations performed as described in the first aspect and any possible method of audio processing described in the first aspect.
[0022] Fourthly, a computer-readable storage medium is provided, the storage medium storing at least one instruction, the instruction being loaded and executed by a processor to perform the operations performed as described in the first aspect and any possible method of audio processing described in the first aspect.
[0023] Fifthly, a computer program product is provided, the computer program product storing at least one instruction, the instruction being loaded and executed by a processor to perform the operations performed as described in the first aspect and any possible implementation of the audio processing method described in the first aspect.
[0024] The beneficial effects of the technical solution provided in this application are: In the solution provided in this application, control audio signals that characterize the chords and rhythm of the original song (first song) are used as the control for music style transfer. The music style characteristics of a reference song (second song) are used as a reference for music style transfer, and text-based music style indication information is used as an auxiliary reference. The music style of the first song is automatically converted to the music style of the second song. The target song audio after music style transfer possesses the music style of the second song, and its chords and rhythm remain consistent with the original song. The entire conversion process requires no manual intervention, saving significant labor costs and achieving high efficiency. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a flowchart of an audio processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of an audio processing method provided in an embodiment of this application; Figure 3 This is a schematic diagram of an audio processing method provided in an embodiment of this application; Figure 4 This is a schematic diagram of an audio processing method provided in an embodiment of this application; Figure 5 This is a schematic diagram of an audio processing method provided in an embodiment of this application; Figure 6 This is a schematic diagram of an audio processing method provided in an embodiment of this application; Figure 7 This is a schematic diagram of an audio processing device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0028] With the continuous advancement of audio technology, users' demands for richer musical styles are also increasing. For example, a song originally stylistically as folk might be replaced by a rock version; this change in musical style is called style transfer. Currently, the common solution for style transfer involves professionals using specialized music editing software to edit the song's audio. However, manual style transfer is costly and inefficient.
[0029] This application provides an audio processing method that can be implemented by a computing device, such as a mobile phone, tablet computer, laptop computer, desktop computer, or a server or server cluster. In this method, a control audio signal of a first song is first acquired, representing the chords and rhythm of the first song. Music style indication information is also acquired, along with music style features of a second song, representing its musical style. Then, the control audio signal of the first song, the music style indication information, and the music style features of the second song are input into a trained music style transfer model to obtain a target song audio. The music style of the target song audio matches the music style of the second song and the music style indicated by the music style indication information. The chords and rhythm of the target song audio also match the chords and rhythm of the first song. In this method, the control audio signal of the original song (first song) is used as the control for audio generation, the musical style characteristics of the reference song (second song) are used as a reference, and the musical style indication information in text form is used as an auxiliary reference. The musical style transfer model is used to convert the musical style of the first song into the musical style of the second song. The target song audio after musical style transfer has the musical style of the second song, and the chords and rhythm are consistent with the original song.
[0030] The audio processing method provided in the embodiments of this application will be described below with reference to the accompanying drawings. Figure 1 The processing flow of this method may include the following steps: Step 101: Obtain the control audio signal of the first song, wherein the control audio signal is used to represent the chords and rhythm of the first song.
[0031] In practice, the terminal device can have a music application installed. If a user wants to change the music style of the first song to that of the second song, they can select the first song as the original song and the second song as the reference song in the aforementioned music application. They can also input music style indication information, which is used to indicate the music style of the audio that the user wants, such as folk, rock, electronic music, etc.
[0032] Furthermore, the terminal device executes the audio processing method provided in the embodiments of this application to achieve music style transfer.
[0033] Alternatively, the terminal device sends a music style migration request to the server, the request carrying the identifier of the original song, the identifier of the reference song, and music style indication information. Then, upon receiving the music style migration request, the server executes the audio processing method provided in this embodiment to achieve music style migration.
[0034] The audio processing method provided in this application embodiment is the same whether it is executed by a terminal or by a server. The following describes the execution of the audio processing method provided in this application embodiment by a computing device. The computing device can be a terminal device or a server, and this application embodiment does not limit it in this regard.
[0035] The computing device acquires the audio of the first song. The audio can be obtained from a music library, provided by the user, or obtained via the Internet. This application embodiment does not limit the method of acquiring the song audio.
[0036] After acquiring the audio of the first song, the computing device can input the audio into the beat recognition model to obtain the beat information of the first song. Then, the beat information of the first song is input into the audio synthesis model to obtain the beat audio signal of the first song. The format of the beat audio signal can be wav (Waveform Audio File Format). The beat recognition model and the audio synthesis model can be pre-trained artificial intelligence models, and in this embodiment, their interfaces can be directly called.
[0037] The computing device inputs the audio of the first song into a chord recognition model to obtain the chord information of the first song. Then, the chord information of the first song is input into an audio synthesis model to obtain the chord audio signal of the first song. The format of the chord audio signal can be WAV, and the chord audio signal can be a bass WAV that can represent chord information. The chord recognition model can be a pre-trained artificial intelligence model, which can be used directly by calling its interface in this embodiment.
[0038] The computing device inputs the audio of the first song into the vocal separation model to obtain the vocal audio signal of the first song. The vocal separation model can be a pre-trained artificial intelligence model, which can be used directly by calling its interface in this embodiment.
[0039] After obtaining the beat audio signal, chord audio signal, and vocal audio signal of the first song, these signals are mixed to obtain the control audio signal of the first song. Specifically, each of the beat audio signal, chord audio signal, and vocal audio signal has a corresponding weight. The computing device can perform a weighted summation of the beat audio signal, chord audio signal, and vocal audio signal according to their respective weights to obtain the control audio signal of the first song. The weights can be configured by relevant technicians according to actual needs, and this application embodiment does not limit this.
[0040] The control audio signal for the first song has the same duration as the audio of the first song.
[0041] Step 102: Obtain text information, which includes music style indication information.
[0042] In practice, the computing device can acquire the aforementioned music style indication information as text information.
[0043] Step 103: Obtain the musical style features of the second song, whereby the musical style features are used to characterize the musical style of the second song.
[0044] In practice, the computing device acquires the audio of the second song. The audio can be obtained from a music library, provided by the user, or obtained via the Internet. This application embodiment does not limit the method of acquiring the song audio.
[0045] After obtaining the audio of the second song, the computing device can input the audio into a style feature extraction model to obtain a music style feature vector of the second song, which serves as the music style feature of the second song. The style feature extraction model can be a pre-trained artificial intelligence model, which can be directly used by calling its interface in this embodiment.
[0046] Step 104: Generate the target song audio based on the control audio signal and text information of the first song and the musical style characteristics of the second song. Specifically, the musical style of the target song audio matches the musical style of the second song and the musical style indicated by the musical style indicator information, and the chords and rhythms of the target song audio match the chords and rhythms of the first song.
[0047] In implementation, the computing device encodes the control audio signal and text information of the first song, as well as the musical style features of the second song, to obtain a first token sequence. Then, the first token sequence is vectorized to obtain a first feature vector, which can be an embedding feature vector. Next, the first feature vector is input into a trained musical style transfer model to obtain the target accompaniment audio. Then, the target accompaniment audio and the vocal audio signal of the first song are mixed to obtain the target song audio. Here, the mixing process can be implemented through a mixing model, which can be a pre-trained artificial intelligence model. In this embodiment, its interface can be directly called.
[0048] In one possible implementation, such as Figure 2As shown, after the computing device performs feature vectorization on the first token sequence, it inputs it into the trained music style transfer model. The music style transfer model outputs a second token sequence. Then, the second token sequence can be input into a scoring model, which outputs a score. If the score is greater than a threshold, the second token sequence is audio decoded to obtain the target accompaniment audio. The scoring model can be a pre-trained artificial intelligence model, which can be directly used by calling its interface in this embodiment.
[0049] In yet another possible implementation, such as Figure 2 As shown, after obtaining the target accompaniment audio, it can be input into a super-resolution model to improve its audio quality, resulting in an enhanced target accompaniment audio. Then, the target accompaniment audio and the vocal audio signal of the first song are mixed to obtain the target song audio. The enhanced target accompaniment audio can be high-definition, stereo, and has a 48kHz sampling rate. High-definition refers to high audio quality standards, such as 24-bit depth. The super-resolution model can be a pre-trained artificial intelligence model; in this embodiment, its interface can be directly called.
[0050] In another possible implementation, to enhance the segmentation of the final target song audio, the target song audio can be generated segment by segment and then combined into a complete target song audio. The song segments can include verses, choruses, pre-choruses, bridges, etc.
[0051] Accordingly, the processing in step 101 above can be as follows: like Figure 3 As shown, after acquiring the audio of the first song, the computing device can input the audio into a segment recognition model. The segment recognition model segments the audio into multiple segments, obtaining segment information for each segment. This segment information indicates whether the audio is a verse, chorus, pre-chorus, or bridge, etc. Then, the computing device can generate the target song accompaniment audio for each segment. The following explanation uses the generation of the target song accompaniment audio for the first segment as an example; the first segment is any one of the multiple segments mentioned above.
[0052] The computing device inputs the first segment of the song's audio into the beat recognition model to obtain the beat information of the first segment. Then, it inputs the beat information and the segment information of the first segment into the audio synthesis model to obtain the beat audio signal of the first segment. Here, the segment information is input to ensure that the beat audio signals of different segments differ in waveform, so that the final generated accompaniment audio has a better sense of segmentation.
[0053] The computing device inputs the first section of the song's audio into the chord recognition model to obtain the chord information for the first section. Then, it inputs the chord information and the segment information of the first section into the audio synthesis model to obtain the chord audio signal for the first section. Here, the segment information is input to ensure that the chord audio signals of different sections differ in waveform, so that the final generated accompaniment audio has a better sense of segmentation.
[0054] The computing device inputs the audio of the first segment of the song into the vocal separation model to obtain the vocal audio signal of the first song.
[0055] After obtaining the beat audio signal, chord audio signal, and vocal audio signal of the first section of the song audio, these signals are mixed to obtain the control audio signal for the first section. Specifically, each of the beat audio signal, chord audio signal, and vocal audio signal has a corresponding weight. The computing device can perform a weighted sum of the beat audio signal, chord audio signal, and vocal audio signal according to their respective weights to obtain the control audio signal for the first section. The control audio signal for the first section has the same duration as the first section of the song audio.
[0056] Furthermore, because segment information is introduced during the generation of control audio signals, the waveforms of the control audio signals for different segments differ. For example, the control audio signal for the verse can consist of one 750Hz and three 350Hz sine waves, the control audio signal for the chorus can consist of alternating 750Hz and 350Hz sine waves, and the control audio signal for the pre-chorus can consist of alternating 750Hz and 150Hz sine waves. In addition, to further emphasize the sense of segmentation, a designated audio signal can be added after the control audio signal for each segment to indicate the segment change; for example, the designated audio signal could be the audio signal for a rise sound effect.
[0057] The processing in step 102 above can be as follows: The computing device can first obtain the lyrics text of the first song and extract the lyrics text of the first paragraph of the first song. The lyrics text of the first paragraph contains the paragraph identifier text of the first paragraph, which is used to indicate which paragraph the first paragraph is.
[0058] Furthermore, the computing device can treat the music style indication information and the lyrics of the first paragraph as text information. With this implementation, when the user inputs the music style indication information, they can input the music style indication information corresponding to each paragraph separately. Therefore, the music style indication information used as text information can refer to the music style indication information corresponding to the first paragraph.
[0059] The processing in step 103 above can be as follows: After obtaining the audio of the second song, the audio can be input into the segment recognition model. The segment recognition model will segment the audio of the second song into segments, and obtain the audio of multiple segments of the second song. The segment information can indicate whether the audio of the song is a verse, chorus, pre-chorus, or bridge, etc.
[0060] In addition, the computing device can input the entire audio of the second song into the style feature extraction model to obtain the music style feature vector of the entire second song.
[0061] Then, the computing device can extract the audio input style features of the second section of the second song using a model to obtain the musical style features of the second section. Furthermore, the computing device can also extract the complete audio input style features of the second song using a model to obtain the musical style features of the entire second song. Specifically, the second section of the second song is of the same type as the first section of the first song. For example, if the first section of the first song is also a verse, then the second section of the second song is a verse; if the first section of the first song is also a chorus, then the second section of the second song is a chorus.
[0062] The processing in step 104 above can be as follows: like Figure 4As shown, the control audio signal and text information corresponding to the first paragraph of the first song, and the audio of the second paragraph of the second song are segmented into words to obtain a third token sequence. Then, a specified candidate token is inserted at the end of the third token sequence to obtain a fourth token sequence. The fourth token sequence is then processed into feature vectors to obtain a second feature vector. Next, the feature sequence corresponding to the specified candidate token in the second feature vector is replaced with the music style feature vector of the entire second song to obtain a third feature vector. Then, the third feature vector is input into the trained music style transfer model to obtain the target song audio of the first paragraph. The music style of the target song audio of the first paragraph matches the music style of the second paragraph of the second song and the music style indicated by the music style indication information. The chords and rhythms of the target song audio of the first paragraph match the chords and rhythms of the first paragraph of the first song. The first paragraph of the first song and the second paragraph of the second song have the same paragraph type. For example, if the first paragraph is a verse, then the second paragraph is also a verse; if the first paragraph is a chorus, then the second paragraph is also a chorus.
[0063] In one possible implementation, the computing device inputs the third feature vector into the trained music style transfer model. The music style transfer model outputs a fifth token sequence. Then, the fifth token sequence can be input into a scoring model, which outputs a score. If the score is greater than a threshold, the fifth token sequence is audio decoded to obtain the target accompaniment audio for the first segment. The same method can be used to obtain the target accompaniment audio for each segment. Then, the accompaniment audio for each target segment is combined according to the playback sequence to obtain the complete target accompaniment audio. Finally, the complete target accompaniment audio is mixed with the vocal audio signal of the first song to obtain the target song audio.
[0064] The music style transfer model described above can be pre-trained. During training, multiple sample songs can be obtained from the music library. For each sample song, the control audio signal of the sample song is extracted as the first input sample, the music style features of the sample song are extracted as the second input sample, and the accompaniment audio of the sample song is extracted as the output sample. In this way, the first input sample, the second input sample, and the output sample corresponding to each sample song can form a sample pair. Each sample pair is used to optimize and fine-tune the music style transfer model until the training termination condition is met, at which point training stops, and the trained music style transfer model is obtained.
[0065] The music style transfer model used above can be an LLM (Large Language Model).
[0066] In addition, such as Figure 5 As shown, after generating the first token sequence, the first token sequence can be delayed first, and then the delayed first token sequence can be vectorized and input into the music style transfer model. Correspondingly, the second token sequence output by the music style transfer model is also a delayed second token sequence. The delayed second token sequence can be de-delayed first, and then audio decoding can be generated to obtain the target accompaniment audio.
[0067] In this embodiment of the application, the input to the music style transfer model has been optimized, such as... Figure 6 As shown, for a multimodal token sequence composed of control audio signals, text information, etc. (such as the third token sequence mentioned above), a designated candidate token is inserted, resulting in a token sequence composed of the multimodal token sequence and the designated candidate token (such as the fourth token sequence mentioned above). Then, the multimodal token sequence and the token sequence composed of the designated candidate token are vectorized to obtain a feature vector, which is also called the latent space of the model input. Furthermore, for the audio of the reference song, its music style features are extracted (such as the music style features of the entire second song mentioned above). Then, these music style features are vectorized to obtain an audio feature vector that matches the latent space. This audio feature vector is then embedded into the latent space of the model input at the position corresponding to the designated candidate token. Finally, the embedded latent space of the model input is input into the music style transfer model.
[0068] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0069] In the solution provided in this application, control audio signals that characterize the chords and rhythm of the original song (the first song) are used as the control for music style transfer. The music style characteristics of a reference song (the second song) are used as the reference for music style transfer, and textual music style indication information is used as an auxiliary reference. A music style transfer model is used to convert the music style of the first song to that of the second song. The target song audio after music style transfer possesses the music style of the second song, and its chords and rhythm remain consistent with the original song. The entire conversion process requires no manual intervention, effectively saving manpower and improving efficiency.
[0070] Based on the same technical concept, embodiments of this application also provide an audio processing apparatus, which can be applied to computing devices, such as... Figure 7 As shown, the device includes an acquisition module 710 and an inference module 720, wherein: The acquisition module 710 is used to extract the beat audio signal, chord audio signal, and vocal audio signal of the first song; mix the beat audio signal, chord audio signal, and vocal audio signal to obtain the control audio signal of the first song, wherein the control audio signal is used to characterize the chords and beat of the first song; acquire text information, wherein the text information includes music style indication information; and acquire the music style features of the second song, wherein the music style features are used to characterize the music style of the second song. The inference module 720 is used to generate song audio based on the control audio signal of the first song, the text information, and the musical style features of the second song to obtain target song audio, wherein the musical style of the target song audio matches the musical style of the second song and the musical style indicated by the musical style indication information, and the chords and rhythms of the target song audio match the chords and rhythms of the first song.
[0071] In one possible implementation, the beat audio signal, chord audio signal, and vocal audio signal are all in the format of waveform audio files (wav).
[0072] In one possible implementation, the acquisition module 710 is used for: Extract the beat audio signal, chord audio signal, and vocal audio signal of the first paragraph of the first song; The beat audio signal, chord audio signal and vocal audio signal are mixed to obtain the control audio signal corresponding to the first paragraph of the first song. The control audio signal corresponding to the first paragraph is used to characterize the chords and beat of the first paragraph, and also to characterize the paragraph information of the first paragraph. In one possible implementation, the acquisition module 710 is used for: Obtain the musical style feature vector of the entire second song and the audio of the second section; The inference module 720 is used for: The control audio signal corresponding to the first paragraph, the text information, and the song audio of the second paragraph are processed by word segmentation to obtain the first word unit token sequence; Insert a specified candidate token at the end of the first token sequence to obtain the second token sequence; The second token sequence is processed into a feature vector to obtain the first feature vector; Replace the feature sequence corresponding to the candidate placeholder in the first feature vector with the music style feature vector of the entire second song to obtain the second feature vector; The second feature vector is input into the trained music style transfer model to obtain the target song audio of the first segment. The music style of the target song audio of the first segment matches the music style of the second segment of the second song and the music style indicated by the music style indication information. The chords and rhythms of the target song audio of the first segment match the chords and rhythms of the first segment of the first song.
[0073] In one possible implementation, the text information also includes the lyrics text of the first paragraph of the first song.
[0074] In one possible implementation, the acquisition module 710 is used for: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the audio of the target song.
[0075] In one possible implementation, the inference module 720 is used for: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the target accompaniment audio. The target accompaniment audio and the vocal audio of the first song are mixed to obtain the target song audio.
[0076] In one possible implementation, the inference module 720 is used for: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the third token sequence. The third token sequence is scored. If the score of the third token sequence is greater than a threshold, the target accompaniment audio is generated based on the third token sequence.
[0077] In the solution provided in this application, control audio signals that characterize the chords and rhythm of the original song (the first song) are used as the control for music style transfer. The music style characteristics of a reference song (the second song) are used as the reference for music style transfer, and textual music style indication information is used as an auxiliary reference. A music style transfer model is used to convert the music style of the first song to that of the second song. The target song audio after music style transfer possesses the music style of the second song, and its chords and rhythm remain consistent with the original song. The entire conversion process requires no manual intervention, effectively saving manpower and improving efficiency.
[0078] It should be noted that the audio processing apparatus provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the terminal device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing apparatus and the audio processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0079] Figure 8 This illustration shows a structural block diagram of a computing device 800 provided in an exemplary embodiment of this application. The computing device 800 may be a terminal device, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The computing device 800 may also be an audio output device, such as headphones or speakers.
[0080] Typically, computing device 800 includes a processor 801 and a memory 802.
[0081] Processor 801 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 801 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 801 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 801 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0082] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 are used to store at least one instruction, which is executed by the processor 801 to implement the audio processing method provided in the method embodiments of this application.
[0083] In some embodiments, the computing device 800 may also optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 803 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, a positioning assembly 808, and a power supply 809.
[0084] Peripheral device interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 801 and memory 802. In some embodiments, processor 801, memory 802 and peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 801, memory 802 and peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0085] The radio frequency (RF) circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 804 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 804 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0086] Display screen 805 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 805 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 801 for processing. In this case, display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 805, disposed on the front panel of computing device 800; in other embodiments, there may be at least two display screens, disposed on different surfaces of computing device 800 or in a folded design; in still other embodiments, display screen 805 may be a flexible display screen, disposed on a curved or folded surface of computing device 800. Furthermore, display screen 805 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 805 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0087] The camera assembly 806 is used to acquire images or videos. Optionally, the camera assembly 806 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0088] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 801 for processing, or input to the radio frequency circuit 804 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the computing device 800. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 807 may also include a headphone jack.
[0089] The positioning component 808 is used to determine the current geographic location of the computing device 800 in order to enable navigation or LBS (Location Based Service). The positioning component 808 can be a positioning component based on GPS (Global Positioning System), BeiDou system, or Galileo system.
[0090] Power supply 809 is used to supply power to the various components in computing device 800. Power supply 809 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When power supply 809 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0091] In some embodiments, the computing device 800 further includes one or more sensors 810. The one or more sensors 810 include, but are not limited to, an accelerometer 811, a gyroscope 812, a pressure sensor 813, a fingerprint sensor 814, an optical sensor 815, and a proximity sensor 816.
[0092] Accelerometer 811 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by computing device 800. For example, accelerometer 811 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 801 can control display screen 805 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 811. Accelerometer 811 can also be used for games or for acquiring user motion data.
[0093] The gyroscope sensor 812 can detect the orientation and rotation angle of the computing device 800. The gyroscope sensor 812, in conjunction with the accelerometer sensor 811, can collect 3D motion data from the user on the computing device 800. Based on the data collected by the gyroscope sensor 812, the processor 801 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0094] The pressure sensor 813 can be disposed on the side bezel of the computing device 800 and / or on the lower layer of the display screen 805. When the pressure sensor 813 is disposed on the side bezel of the computing device 800, it can detect the user's grip signal on the computing device 800, and the processor 801 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 813. When the pressure sensor 813 is disposed on the lower layer of the display screen 805, the processor 801 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 805. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0095] The fingerprint sensor 814 is used to collect a user's fingerprint. The processor 801 identifies the user based on the fingerprint collected by the fingerprint sensor 814, or vice versa. When the user's identity is verified as trusted, the processor 801 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 814 can be located on the front, back, or side of the computing device 800. When the computing device 800 has physical buttons or a manufacturer's logo, the fingerprint sensor 814 can be integrated with the physical buttons or manufacturer's logo.
[0096] An optical sensor 815 is used to collect ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 based on the ambient light intensity collected by the optical sensor 815. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 based on the ambient light intensity collected by the optical sensor 815.
[0097] A proximity sensor 816, also known as a distance sensor, is typically located on the front panel of the computing device 800. The proximity sensor 816 is used to detect the distance between the user and the front of the computing device 800. In one embodiment, when the proximity sensor 816 detects that the distance between the user and the front of the computing device 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from a screen-on state to a screen-off state; when the proximity sensor 816 detects that the distance between the user and the front of the computing device 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from a screen-off state to a screen-on state.
[0098] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the computing device 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0099] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to perform the audio processing method described above. This computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices, etc.
[0100] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices) involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, all audio and text data involved in this application were obtained with full authorization.
[0101] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0102] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0103] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An audio processing method, characterized in that, The method includes: Extract the beat audio signal, chord audio signal, and vocal audio signal from the first song's audio. The beat audio signal, chord audio signal and vocal audio signal are mixed to obtain the control audio signal of the first song, wherein the control audio signal is used to characterize the chords and beat of the first song; Obtain text information, wherein the text information includes music style indication information; Obtain the musical style features of the second song, wherein the musical style features are used to characterize the musical style of the second song; Based on the control audio signal of the first song, the text information, and the musical style features of the second song, a target song audio is generated. The musical style of the target song audio matches the musical style of the second song and the musical style indicated by the musical style indication information. The chords and rhythms of the target song audio match the chords and rhythms of the first song.
2. The method according to claim 1, characterized in that, The extraction of the beat audio signal, chord audio signal, and vocal audio signal of the first song includes: Extract the beat audio signal, chord audio signal, and vocal audio signal of the first paragraph of the first song; The control audio signal for the first song obtained by mixing the beat audio signal, chord audio signal, and vocal audio signal includes: The beat audio signal, chord audio signal and vocal audio signal are mixed to obtain the control audio signal corresponding to the first paragraph of the first song. The control audio signal corresponding to the first paragraph is used to represent the chords and beat of the first paragraph, and also to represent the paragraph information of the first paragraph.
3. The method according to claim 2, characterized in that, The acquisition of the musical style characteristics of the second song includes: Obtain the musical style feature vector of the entire second song and the audio of the second section; The process of generating the target song audio based on the control audio signal of the first song, the text information, and the musical style features of the second song includes: The control audio signal corresponding to the first paragraph, the text information, and the song audio of the second paragraph are processed by word segmentation to obtain the first word unit token sequence; Insert a specified candidate token at the end of the first token sequence to obtain the second token sequence; The second token sequence is processed into a feature vector to obtain the first feature vector; Replace the feature sequence corresponding to the specified candidate token in the first feature vector with the music style feature vector of the entire second song to obtain the second feature vector; The second feature vector is input into the trained music style transfer model to obtain the target song audio of the first segment. The music style of the target song audio of the first segment matches the music style of the second segment of the second song and the music style indicated by the music style indication information. The chords and rhythms of the target song audio of the first segment match the chords and rhythms of the first segment of the first song.
4. The method according to claim 3, characterized in that, The text information also includes the lyrics of the first paragraph of the first song.
5. The method according to claim 1, characterized in that, The process of generating the target song audio based on the control audio signal of the first song, the text information, and the musical style features of the second song includes: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the audio of the target song.
6. The method according to claim 5, characterized in that, The step of inputting the control audio signal of the first song, the text information, and the musical style features of the second song into the trained musical style transfer model to obtain the target song audio includes: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the target accompaniment audio. The target accompaniment audio and the vocal audio of the first song are mixed to obtain the target song audio.
7. The method according to claim 6, characterized in that, The step of inputting the control audio signal of the first song, the text information, and the musical style features of the second song into the trained musical style transfer model to obtain the target accompaniment audio includes: The control audio signal of the first song, the text information, and the musical style features of the second song are input into the trained musical style transfer model to obtain the third token sequence. The third token sequence is scored. If the score of the third token sequence is greater than a threshold, the target accompaniment audio is generated based on the third token sequence.
8. A computing device, characterized in that, The computing device includes a processor and a memory, the memory storing at least one instruction that is loaded and executed by the processor to perform the operations performed in the audio processing method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to perform the operation of the audio processing method as described in any one of claims 1-7.
10. A computer program product, characterized in that, The computer program product stores at least one instruction, which is loaded and executed by a processor to perform the operation performed by the audio processing method as described in any one of claims 1-7.