Music conversion method and related device, equipment and storage medium

By extracting and constructing music feature pairs, automatic conversion of music styles is achieved, which solves the problems of high cost and poor effect in existing technologies and improves the accuracy and granularity of music conversion.

CN119811403BActive Publication Date: 2025-10-17IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411802563.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-10-17
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

In the existing technology, music style conversion requires high costs of professional mixers or bands and has poor results, which cannot meet users' needs for different music styles.

Method used

By extracting the audio features and sub-features of the music to be converted and the reference music, constructing sub-feature pairs, and generating the target music, the automatic conversion of music styles is achieved.

Benefits of technology

Improves the accuracy and granularity of music conversion and reduces the loss of detail data in complex audio signals during the conversion process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811403B_ABST
    Figure CN119811403B_ABST
Patent Text Reader

Abstract

The application discloses a music conversion method and related devices, equipment and storage media, wherein the music conversion method comprises: based on the audio data of the music to be converted, extracting the first audio features of the music to be converted and the first sub-features about several music components, and based on the audio data of the reference music, extracting the second audio features of the reference music and the second sub-features about several music components; based on the music components of each first sub-feature and the music components of each second sub-feature, constructing sub-feature pairs belonging to different music components; wherein the sub-feature pair of any music component represents the characteristic mapping relationship of the second sub-feature belonging to the music component about the music to be converted; based on the sub-feature pair, the first audio feature and the second audio feature, generating target music. The above scheme can realize automatic conversion of music style and improve the accuracy of music conversion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio processing, in particular to a music conversion method and related device, equipment and storage medium. BACKGROUND

[0002] Due to the increasingly diversified needs of users for music, different users expect to choose a performance version that meets their favorite style for the same song.

[0003] In the prior art, a song is usually re-performed based on professional sound mixers or bands, or the spectrum of the music to be converted is migrated to achieve music style conversion. However, professional sound mixers or bands require high performance costs, and due to the complexity of music signals, the target music effect obtained after spectrum migration is poor, which cannot meet the needs of users for different music styles. Therefore, how to realize automatic conversion of music style and improve the accuracy of music conversion has become a problem to be solved. SUMMARY

[0004] The technical problem solved by the present application is to provide a music conversion method and related device, equipment and storage medium, which can realize automatic conversion of music style and improve the accuracy of music conversion.

[0005] To solve the above technical problem, the first aspect of the present application provides a music conversion method, comprising: based on the audio data of the music to be converted, extracting the first audio feature of the music to be converted and the first sub-feature about several music components, and based on the audio data of the reference music, extracting the second audio feature of the reference music and the second sub-feature about several music components; based on the music components of each first sub-feature and the music components of each second sub-feature, constructing a sub-feature pair belonging to different music components; wherein any music component sub-feature pair represents: the feature mapping relationship of the second sub-feature belonging to the music component about the music to be converted; based on the sub-feature pair, the first audio feature and the second audio feature, generating target music.

[0006] To solve the above technical problems, the second aspect of the present application provides a music conversion device, comprising: an extraction module, a construction module and a generation module, the extraction module is used for extracting first audio features of the music to be converted and first sub-features about several music components based on audio data of the music to be converted, and extracting second audio features of the reference music and second sub-features about several music components based on audio data of the reference music; the construction module is used for constructing sub-feature pairs belonging to different music components based on music components of each first sub-feature and music components of each second sub-feature; wherein the sub-feature pair of any music component represents the characteristic mapping relationship of the second sub-feature belonging to the music component with respect to the music to be converted; the generation module is used for generating target music based on the sub-feature pair, the first audio feature and the second audio feature.

[0007] To solve the above technical problems, the third aspect of the present application provides an electronic device, comprising a memory and a processor coupled with each other, the memory stores program instructions, and the processor is used to execute the program instructions to realize the music conversion method in the first aspect.

[0008] To solve the above technical problems, the fourth aspect of the present application provides a computer readable storage medium, which stores program instructions capable of being executed by a processor, and the program instructions are used to realize the music conversion method in the first aspect.

[0009] The above scheme extracts first audio features of the music to be converted and first sub-features about several music components based on audio data of the music to be converted, and extracts second audio features of the reference music and second sub-features about several music components based on audio data of the reference music, constructs sub-feature pairs belonging to different music components based on music components of each first sub-feature and music components of each second sub-feature, the sub-feature pair of any music component represents the characteristic mapping relationship of the second sub-feature belonging to the music component with respect to the music to be converted, and generates target music based on the sub-feature pair, the first audio feature and the second audio feature. Therefore, based on the audio data of the music to be converted and the audio data of the reference music, the music of the desired style can be automatically converted, and the sub-feature pairs belonging to different music components are used as feature data for generating target music, which can preserve the feature matching relationship between the music to be converted and the reference music at the music component level as much as possible, improve the granularity of target music generation, and reduce the loss of detailed data in the conversion process as much as possible for music with complex audio signals. Therefore, the automatic conversion of music style can be realized, and the accuracy of music conversion can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 is a flowchart of an embodiment of the music conversion method of the present application;

[0011] Figure 2 is a framework schematic diagram of an embodiment of the music conversion model in the music conversion method of the present application;

[0012] Figure 3 is a framework schematic diagram of an embodiment of the music conversion device of the present application;

[0013] Figure 4 is a framework schematic diagram of an embodiment of the electronic device of the present application;

[0014] Figure 5 is a framework schematic diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION

[0015] The scheme of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0016] In the following description, specific details such as specific system structures, interfaces, techniques, etc. are presented in order to provide a thorough understanding of the present application for the sake of explanation, but not for the sake of limitation.

[0017] If the technical scheme of the present application involves personal information, the product applying the technical scheme of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical scheme of the present application involves sensitive personal information, the product applying the technical scheme of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent sign is set to inform that the personal information collection range has been entered, and the personal information will be collected. If the individual voluntarily enters the collection range, it is considered to agree to collect the personal information. Or, in the case of informing the personal information processing rules through obvious signs / information on the device for processing personal information, the personal authorization is obtained through the pop-up information or the request of the individual to upload his / her personal information, etc. The personal information processing rules can include personal information processor, personal information processing purpose, processing method, and personal information type, etc.

[0018] The terms "system" and "network" are often used interchangeably in this paper. The term "and / or" in this paper is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the segment " / " in this paper generally represents that the front and rear associated objects are in an "or" relationship. In addition, "multiple" in this paper means two or more than two.

[0019] Please refer to Figure 1 , Figure 1 is a flow schematic diagram of an embodiment of the music conversion method of the present application. Specifically, it can include the following steps:

[0020] Step S10: based on the audio data of the music to be converted, extracting the first audio features of the music to be converted and the first sub-features about several music components, and based on the audio data of the reference music, extracting the second audio features of the reference music and the second sub-features about several music components.

[0021] In the embodiments of the present disclosure, the specific types of the music to be converted and the reference music are not limited in the present application, for example, music composed of multi-instrument audio and single-person audio, music composed of multi-instrument audio, music composed of multi-person audio, etc. Specifically, the music to be converted can be a folk-style song A, and the reference music can be a rock-style song B.

[0022] In one implementation scenario, the first audio features of the music to be converted can reflect the overall properties of the music to be converted, such as the style, rhythm, audio content, etc. of the music to be converted, and the second audio features of the reference music can reflect the overall properties of the reference music, such as the style, rhythm, audio content, etc. of the reference music.

[0023] In one specific implementation scenario, the first audio features and the second audio features can be implemented based on mel-frequency cepstral coefficients, spectrograms, mel-frequency spectra, etc. For example, for each frame of audio data, MFCC features are calculated. MFCC provides an effective audio feature representation by simulating the hearing characteristics of the human ear. The calculation process includes: applying a discrete Fourier transform (DFT) to obtain the frequency spectrum of the frame, mapping the frequency spectrum to a mel scale, and using a set of mel filters. Taking the logarithm of the output of each mel filter, applying a discrete cosine transform (DCT) to the logarithmic mel spectrum to obtain MFCC features as the audio features of the target speech segment. It should be noted that the extraction method of the first audio features and the second audio features is not limited in the present application, and the technical details can be referred to the above technical method. For the sake of brevity, it will not be repeated here.

[0024] In one implementation scenario, a music component represents an audio component of music, such as drums, bass, piano, vocals, electronic mixing, etc. The timbre, loudness, frequency, etc. of each music component are different, the melody, content, etc. of the same music component in different music are not the same, and the types of music components contained in music of different styles are also not the same, for example, the folk-style song A contains guitar, bass, hand drum, harmonica, sand hammer, vocals, etc., and the rock-style song B contains electronic guitar, bass, drums, electronic mixing, etc.

[0025] In one specific implementation scenario, the target sub-audio of the music components in the target music is obtained, and the target sub-feature is obtained based on the target sub-audio. Specifically, in the case of the target music being the to-be-converted music, the target sub-feature is the first sub-feature; and in the case of the target music being the reference music, the target sub-feature is the second sub-feature.

[0026] In one specific implementation scenario, the target sub-audio of each type of music component in the target music is obtained based on the audio information of the target music. For example, the target sub-audio belonging to each instrument or the target sub-audio belonging to each singer is obtained based on the tone, loudness, frequency, etc. in the audio information. As one possible implementation, a music component separation model can be pre-trained. The music component separation model can include, but is not limited to, an Encoder-Decoder architecture network model, etc. In order to ensure the generation accuracy of the music component separation model as much as possible, sample music composed of multiple sample sub-audios can be collected, and the sample sub-audios are labeled with real music components. Based on this, the sample music can be processed based on the music component separation model to obtain predicted sub-audios of the sample music and predicted music components of each predicted sub-audio. Thus, the network parameters of the music component separation model can be adjusted based on the difference between the sample sub-audios and the predicted sub-audios, and the difference between the real music components and the predicted music components, until the music component separation model converges. The target sub-audio of each music component can be obtained based on the music component separation model that has converged.

[0027] In another specific implementation scenario, the target sub-feature is obtained by encoding the target music by a sub-component encoder. The specific network structure of the sub-component encoder is not limited in the present application, such as a conditional similarity network structure, etc.

[0028] In a specific implementation scenario, the training process of the sub-component encoder includes: obtaining sample music containing a plurality of music component sample sub-audios, and each sample sub-audio is labeled with a sample music component; obtaining positive example sample audios and negative example sample audios about the sample music component, the positive example sample audios are consistent with the sample music component of the sample audio, and the negative example sample audios are inconsistent with the sample music component of the sample audio, for example, the positive example sample audios are audios of the same type of musical instruments in different songs, and the negative example sample audios are audios of different types of musical instruments in the same song; respectively encoding the sample sub-audio, the positive example sample audio, and the negative example sample audio about the sample music component based on the sub-component encoder to obtain sample sub-audio features, positive example sample audio features, and negative example sample audio features; and adjusting network parameters of the sub-component encoder based on a difference between the sample sub-audio features, the positive example sample audio features, and the negative example sample audio features measured by a triplet loss function. The above scheme trains the sub-component encoder through the triplet loss function, optimizes the sub-component encoder by defining a triplet, makes sample sub-audios of the same music component closer in the feature space, and makes sample sub-audios of different music components farther away from each other, thereby improving the accuracy of feature encoding of target sub-audios of different music components and improving the effect of music conversion.

[0029] It should be noted that the extraction method of the first sub-feature and the extraction method of the second sub-feature can be the same or different, which is not limited in the present application.

[0030] In an implementation scenario, before extracting the first audio feature of the music to be converted and the first sub-feature about a plurality of music components based on the audio data of the music to be converted, a standard instrument encoding library is obtained, the standard instrument encoding library contains standard encodings of each instrument about a basic pitch, and the basic pitch, also known as the fundamental pitch, refers to seven pitches with independent names in the musical sound system, which are consistent with the sound emitted by the white keys on the piano. The basic pitch is represented by seven letters "C, D, E, F, G, A, B" or "do, re, mi, fa, sol, la, si", which are called "note names". It can be understood that the audio information of different instruments or different singers about the basic pitch may be different, specifically the differences in loudness, tone color, frequency, etc., and the audio information of the same type of instrument about the basic pitch is relatively similar. Therefore, each standard encoding in the standard instrument encoding library can reflect the basic audio information of a type of instrument and is not affected by melody, rhythm, etc.

[0031] It should be noted that the method of obtaining the standard encoding of each instrument about the basic pitch is not limited in the present application.

[0032] In a specific implementation scenario, in a case where the music component of the first sub-feature belongs to a musical instrument, after extracting a plurality of first sub-features about the music to be converted, and before constructing the pair of sub-features belonging to different music components, a first standard code consistent with the music component of the first sub-feature is selected from a standard musical instrument code library, and a fusion code obtained by fusing the first sub-feature and the first standard code is taken as an updated first sub-feature. The above scheme embeds the standard code about the musical instrument component into the first sub-feature of the music to be converted to obtain the updated first sub-feature, which can improve the specificity of the sub-feature of different music components, and can also reduce the loss of detailed data as much as possible in the conversion process for music with complex audio signals, and thus can further improve the accuracy of music conversion.

[0033] In a specific implementation scenario, the fusion of the first sub-feature and the first standard code is implemented by feature splicing or the like to obtain the updated first sub-feature.

[0034] In a specific implementation scenario, in a case where the music component of the second sub-feature belongs to a musical instrument, after extracting a plurality of second sub-features about the reference music, and before constructing the pair of sub-features belonging to different music components, a second standard code consistent with the music component of the second sub-feature is selected from a standard musical instrument code library, and a fusion code obtained by fusing the second sub-feature and the second standard code is taken as an updated second sub-feature. The above scheme embeds the standard code about the musical instrument component into the second sub-feature of the music to be converted to obtain the updated second sub-feature, which can improve the specificity of the sub-feature of different music components, and can also reduce the loss of detailed data as much as possible in the conversion process for music with complex audio signals, and thus can further improve the accuracy of music conversion.

[0035] In a specific implementation scenario, the fusion of the second sub-feature and the second standard code is implemented by feature splicing or the like to obtain the updated second sub-feature.

[0036] Step S20: constructing a pair of sub-features belonging to different music components based on the music component of each first sub-feature and the music component of each second sub-feature.

[0037] In the embodiments of the present disclosure, the pair of sub-features of any music component represents a feature mapping relationship of the second sub-feature belonging to the music component about the music to be converted. For example, the feature mapping relationship of the second sub-feature belonging to the guitar about the music to be converted, and the pair of sub-features can represent the sub-feature conversion logic at the music component level.

[0038] In an implementation scenario, in a case where there is a first sub-feature consistent with the music component of the second sub-feature, a sub-feature pair is constructed based on the second sub-feature, the first sub-feature consistent with the music component of the second sub-feature. For example, there is a second sub-feature A with a music component of guitar, and there is a first sub-feature B with a music component of guitar, and a sub-feature pair (A, B) of guitar is constructed.

[0039] In an implementation scenario, in a case where there is no first sub-feature consistent with the music component of the second sub-feature, a sub-feature pair is constructed based on the second sub-feature and a zero vector feature. For example, there is a second sub-feature C with a music component of piano, but there is no first sub-feature with a music component of piano, and a zero feature vector is set as M, and a sub-feature pair (C, M) of piano is constructed.

[0040] In a specific implementation scenario, when there is no second sub-feature consistent with the music component of the first sub-feature, no sub-feature pair is constructed for the music component. The above scheme can generate a sub-feature pair based on the music component attribute of the reference music as much as possible, and improve the accuracy of target music generation.

[0041] Step S30: generating target music based on the sub-feature pair, the first audio feature, and the second audio feature.

[0042] In an implementation scenario, based on the audio data of the music to be converted and the audio data of the reference music, automatic conversion of music of a desired style can be realized, and the sub-feature pair belonging to different music components is used as feature data for generating target music, which can preserve the feature matching relationship between the music to be converted and the reference music at the music component level as much as possible, improve the granularity of target music generation, and reduce the loss of detailed data in the conversion process as much as possible for music with complex audio signals. Therefore, automatic conversion of music style can be realized, and the accuracy of music conversion can be improved.

[0043] In an implementation scenario, based on each sub-feature pair, the first audio feature, and the second audio feature, an audio hidden layer feature of each sub-feature pair belonging to each music component is generated, specifically, the audio hidden layer feature of the sub-feature pair is a vector, a matrix, etc., which is not limited in the present application. Based on each audio hidden layer feature, the first audio feature, and the second audio feature, target music is decoded. The above scheme can represent the encoding features of each music component with respect to the music to be converted and the reference music through the audio hidden layer feature of the sub-feature pair, which can preserve the feature matching relationship between the music to be converted and the reference music at the music component level as much as possible, improve the granularity of target music generation, and reduce the loss of detailed data in the conversion process as much as possible for music with complex audio signals. Therefore, based on each audio hidden layer feature, the first audio feature, and the second audio feature, the target music generated can improve the accuracy of music generation.

[0044] In one specific implementation scenario, before generating the audio latent feature of the sub-feature pair belonging to each music component based on each sub-feature pair, the first audio feature and the second audio feature, respectively, the first weight, the second weight and the third weight are obtained based on the music component to which the sub-feature pair belongs, the first weight is not less than the sum of the numerical values of the second weight and the third weight, the first weight represents the importance of the sub-feature pair in the audio latent feature generation stage, the second weight represents the importance of the first audio feature in the audio latent feature generation stage, and the third weight represents the importance of the second audio feature in the audio latent feature generation stage. The above scheme, the first weight is not less than the sum of the numerical values of the second weight and the third weight, which represents that more emphasis is placed on the importance of each music component sub-feature pair in the audio latent feature generation stage, improving the granularity of target music generation, and reducing the loss of detailed data in the conversion process as much as possible for music with complex audio signals.

[0045] It should be noted that the method of obtaining the weight is not limited in the present application, for example, a fixed value set by the technician, a dynamic value dynamically adjusted based on the music type and the music component type, etc.

[0046] In one specific implementation scenario, based on the first weight, the second weight and the third weight corresponding to the music component to which the sub-feature pair belongs, the sub-feature pair, the first audio feature and the second audio feature are weighted respectively to obtain the weighted sub-feature pair, the first audio feature and the second audio feature, and the audio latent feature of the sub-feature pair is obtained based on the weighted sub-feature pair, the first audio feature and the second audio feature. The above scheme, the audio latent feature of the sub-feature pair can represent the encoding features of each music component with respect to the to-be-converted music and the reference music, and can preserve the feature matching relationship between the to-be-converted music and the reference music at the music component level as much as possible, improving the granularity of target music generation, and reducing the loss of detailed data in the conversion process as much as possible for music with complex audio signals. Therefore, the target music generated based on each audio latent feature, the first audio feature and the second audio feature can improve the accuracy of music generation.

[0047] The scheme is based on audio data of the music to be converted, extracts first audio features of the music to be converted and first sub-features about several music components, and based on audio data of the reference music, extracts second audio features of the reference music and second sub-features about several music components, constructs sub-feature pairs belonging to different music components based on music components of each first sub-feature and music components of each second sub-feature, and any music component sub-feature pair represents: the second sub-feature belonging to the music component is about the feature mapping relationship of the music to be converted, and based on the sub-feature pair, the first audio feature and the second audio feature, generates the target music. Therefore, based on the audio data of the music to be converted and the audio data of the reference music, the music of the desired style can be automatically converted, and the sub-feature pairs belonging to different music components are used as feature data for generating the target music, which can preserve the feature matching relationship between the music to be converted and the reference music at the music component level as much as possible, improve the granularity of the target music generation, and reduce the loss of detailed data in the conversion process as much as possible for music with complex audio signals. Therefore, the automatic conversion of music style can be realized, and the accuracy of music conversion can be improved.

[0048] Please refer to Figure 2 , Figure 2 is a schematic diagram of an embodiment of the music conversion model 20 in the music conversion method of the present application. In a specific implementation scenario, the target music is generated by the music conversion model 20 based on the input music to be converted and the reference music, as shown in Figure 2 The music conversion model 20 includes a sub-component encoding network 211, a feature fusion network 220, a sub-feature pair construction network 230, a sub-component decoding network 240, a global decoding network 250, and a feature extraction network 212. The sub-component encoding network 211 is used to extract first sub-features of the music to be converted about several music components and second sub-features of the reference music about several music components. The feature extraction network 212 is used to extract first audio features of the music to be converted and second audio features of the reference music. The feature fusion network 220 is used to fuse the first sub-features with the first standard encoding to obtain updated first sub-features, and fuse the second sub-features with the second standard encoding to obtain updated second sub-features. The sub-feature pair construction network 230 is used to construct sub-feature pairs of different music components based on the first sub-features and the second sub-features. The sub-component decoding network 240 is used to generate audio hidden layer features of the sub-feature pairs belonging to each music component based on the sub-feature pairs, the first audio features, and the second audio features. The global decoding network 250 is used to decode to obtain the target music based on each audio hidden layer feature, the first audio feature, and the second audio feature.

[0049] In a specific implementation scenario, the sub-component decoding network 240 is a shallow decoder, and focuses on the music component migration task. The sub-component decoding network 240 weights the sub-feature pair, the first audio feature, and the second audio feature based on the first weight, the second weight, and the third weight corresponding to the music component to which the sub-feature pair belongs, respectively, to obtain the weighted sub-feature pair, the first audio feature, and the second audio feature. Based on the weighted sub-feature pair, the first audio feature, and the second audio feature, the audio hidden layer feature of the sub-feature pair is processed. Specifically, the weighting processing can be implemented through a gating weighting structure. The global decoding network 250 focuses on the overall style migration task, can constrain the range of data processing effect, and improves the accuracy of the target music generation.

[0050] It should be noted that the specific network structures of the sub-component encoding network 211, the feature fusion network 220, the sub-feature pair construction network 230, the sub-component decoding network 240, the global decoding network 250, and the feature extraction network 212 are not limited in the present application. For brevity, the specific implementation methods are not described here.

[0051] The above scheme, the music conversion model 20 extracts the first audio feature of the to-be-converted music and the first sub-feature of the plurality of music components based on the audio data of the to-be-converted music, extracts the second audio feature of the reference music and the second sub-feature of the plurality of music components based on the audio data of the reference music, constructs the sub-feature pair belonging to different music components based on the music component of each first sub-feature and the music component of each second sub-feature, and the sub-feature pair of any music component represents the feature mapping relationship of the second sub-feature belonging to the music component with respect to the to-be-converted music. Based on the sub-feature pair, the first audio feature, and the second audio feature, the target music is generated. Therefore, based on the audio data of the to-be-converted music and the audio data of the reference music, the music of the desired style can be automatically converted, and the sub-feature pair belonging to different music components is used as the feature data for generating the target music, which can preserve the feature matching relationship between the to-be-converted music and the reference music at the music component level as much as possible, improve the fine granularity of the target music generation, and reduce the loss of detailed data in the conversion process as much as possible for music with complex audio signals. Therefore, the automatic conversion of music style can be realized, and the accuracy of music conversion can be improved.

[0052] Please refer to Figure 3 , Figure 3 is a frame schematic diagram of an embodiment of the music conversion device 30 of the present application. As Figure 3As shown, the music conversion device 30 comprises an extraction module 31, a construction module 32 and a generation module 33. The extraction module 31 is configured to extract, based on audio data of the music to be converted, first audio features of the music to be converted and first sub-features of a plurality of music components, and extract, based on audio data of the reference music, second audio features of the reference music and second sub-features of the plurality of music components. The construction module 32 is configured to construct, based on the music components of each first sub-feature and the music components of each second sub-feature, sub-feature pairs belonging to different music components. Any sub-feature pair of a music component represents a characteristic mapping relationship of the second sub-feature belonging to the music component with respect to the music to be converted. The generation module 33 is configured to generate, based on the sub-feature pairs, the first audio features and the second audio features, target music.

[0053] According to the above scheme, the music conversion device 30 extracts, based on audio data of the music to be converted, first audio features of the music to be converted and first sub-features of a plurality of music components, and extracts, based on audio data of the reference music, second audio features of the reference music and second sub-features of the plurality of music components. The music conversion device 30 constructs, based on the music components of each first sub-feature and the music components of each second sub-feature, sub-feature pairs belonging to different music components. Any sub-feature pair of a music component represents a characteristic mapping relationship of the second sub-feature belonging to the music component with respect to the music to be converted. The music conversion device 30 generates, based on the sub-feature pairs, the first audio features and the second audio features, target music. Therefore, based on the audio data of the music to be converted and the audio data of the reference music, the music conversion device 30 can automatically convert the music of the desired style. By taking the sub-feature pairs belonging to different music components as feature data for generating the target music, the music conversion device 30 can preserve the feature matching relationship between the music to be converted and the reference music at the music component level as much as possible, improve the granularity of the generated target music, and reduce the loss of detailed data in the conversion process as much as possible for music with complex audio signals. Therefore, the music conversion device 30 can realize automatic conversion of music style and improve the accuracy of music conversion.

[0054] In some disclosed embodiments, the generation module 33 further comprises a hidden layer feature generation module (not shown) configured to generate, based on each sub-feature pair, the first audio features and the second audio features, audio hidden layer features of the sub-feature pair belonging to each music component. The generation module 33 further comprises an audio feature decoding module (not shown) configured to decode, based on each audio hidden layer feature, the first audio features and the second audio features, the target music.

[0055] In some disclosed embodiments, before generating the audio latent feature of the sub-feature pair belonging to each music component based on each sub-feature pair, the first audio feature and the second audio feature, the music conversion device 30 further comprises a weight obtaining module (not shown) for obtaining a first weight, a second weight and a third weight based on the music component to which the sub-feature pair belongs; wherein the first weight is not less than the sum of the numerical values of the second weight and the third weight, the first weight represents the importance of the sub-feature pair in the audio latent feature generation stage, the second weight represents the importance of the first audio feature in the audio latent feature generation stage, and the third weight represents the importance of the second audio feature in the audio latent feature generation stage; the latent feature generation module (not shown) further comprises a weighting module (not shown) for weighting the sub-feature pair, the first audio feature and the second audio feature based on the first weight, the second weight and the third weight corresponding to the music component to which the sub-feature pair belongs, to obtain the weighted sub-feature pair, the first audio feature and the second audio feature; the latent feature generation module (not shown) further comprises a processing module (not shown) for processing the audio latent feature of the sub-feature pair based on the weighted sub-feature pair, the first audio feature and the second audio feature.

[0056] In some disclosed embodiments, the construction module 32 further comprises a first construction module (not shown) for constructing a sub-feature pair based on the second sub-feature and the first sub-feature consistent with the music component of the second sub-feature in response to the existence of the first sub-feature consistent with the music component of the second sub-feature; the construction module 32 further comprises a second construction module (not shown) for constructing a sub-feature pair based on the second sub-feature and the zero vector feature in response to the non-existence of the first sub-feature consistent with the music component of the second sub-feature.

[0057] In some disclosed embodiments, the music conversion device further comprises a sub-audio obtaining module (not shown) for obtaining target sub-audio of several music components in the target music; the music conversion device further comprises a sub-feature encoding module (not shown) for encoding to obtain target sub-features based on the target sub-audio; wherein, in the case that the target music is the music to be converted, the target sub-features are the first sub-features, and in the case that the target music is the reference music, the target sub-features are the second sub-features.

[0058] In some disclosed embodiments, the target sub-feature is obtained by encoding the target music by a sub-component encoder, and the music conversion device 30 further comprises a sub-component encoder training module (not shown) configured to obtain sample music comprising a plurality of sample sub-audios of music components; each sample sub-audio is labeled with a sample music component; obtain positive example sample audios and negative example sample audios of the sample music component; the positive example sample audios are consistent with the sample music component of the sample audio, and the negative example sample audios are inconsistent with the sample music component of the sample audio; encode the sample sub-audio, the positive example sample audio, and the negative example sample audio of the sample music component by the sub-component encoder respectively to obtain sample sub-audio features, positive example sample audio features, and negative example sample audio features; and adjust network parameters of the sub-component encoder based on a difference between the sample sub-audio features, the positive example sample audio features, and the negative example sample audio features measured by a triplet loss function.

[0059] In some disclosed embodiments, before extracting the first audio features of the music to be converted and the first sub-features of the plurality of music components based on audio data of the music to be converted, the music conversion device 30 further comprises an encoding library obtaining module (not shown) configured to obtain a standard instrument encoding library; the standard instrument encoding library comprises standard encodings of each instrument with respect to a basic pitch; in the case that the music component of the first sub-feature belongs to an instrument, after extracting the plurality of first sub-features of the music to be converted, and before constructing the pair of sub-features belonging to different music components, the music conversion device 30 further comprises a first updating module (not shown) configured to select a first standard encoding in the standard instrument encoding library consistent with the music component of the first sub-feature; and fuse the first sub-feature and the first standard encoding to obtain a fusion encoding as the updated first sub-feature; in the case that the music component of the second sub-feature belongs to an instrument, after extracting the plurality of second sub-features of the reference music, and before constructing the pair of sub-features belonging to different music components, the music conversion device 30 further comprises a second updating module (not shown) configured to select a second standard encoding in the standard instrument encoding library consistent with the music component of the second sub-feature; and fuse the second sub-feature and the second standard encoding to obtain a fusion encoding as the updated second sub-feature.

[0060] Please refer to Figure 4 , Figure 4 is a schematic diagram of a framework of an embodiment of the electronic device 40. The electronic device 40 comprises a memory 41 and a processor 42, the memory 41 stores program instructions, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above music conversion method embodiments. For details, please refer to the foregoing disclosed embodiments, which will not be repeated here. The electronic device 40 can specifically include but is not limited to: a server, a smart phone, a notebook computer, a tablet computer, a self-service machine, etc., which are not limited here.

[0061] In particular, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the above-mentioned embodiments of the music conversion method. The processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 can be an integrated circuit chip having a processing capability of signals. The processor 42 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 42 can be implemented by an integrated circuit chip together.

[0062] According to the above-mentioned scheme, the electronic device 40 extracts the first audio features of the music to be converted and the first sub-features of the music components based on the audio data of the music to be converted, extracts the second audio features of the reference music and the second sub-features of the music components based on the audio data of the reference music, constructs the sub-feature pairs belonging to different music components based on the music components of each first sub-feature and the music components of each second sub-feature, and the sub-feature pair of any music component represents the feature mapping relationship of the second sub-feature belonging to the music component with respect to the music to be converted, and generates the target music based on the sub-feature pairs, the first audio features and the second audio features. Therefore, based on the audio data of the music to be converted and the audio data of the reference music, the music conversion of the desired style can be automatically realized, and the sub-feature pairs belonging to different music components are used as the feature data for generating the target music, which can preserve the feature matching relationship between the music to be converted and the reference music at the music component level as much as possible, improves the granularity of the target music generation, and reduces the loss of detailed data in the conversion process as much as possible for the music with complex audio signals. Therefore, the automatic conversion of the music style can be realized, and the accuracy of the music conversion can be improved.

[0063] Please refer to Figure 5 , Figure 5 is a schematic diagram of an embodiment of the computer readable storage medium 50 of the present application. The computer readable storage medium 50 stores program instructions 51 capable of being executed by the processor, and the program instructions 51 are used to implement the steps in any of the above-mentioned embodiments of the music conversion method.

[0064] The scheme extracts the first audio features of the music to be converted and the first sub-features of the music components based on the audio data of the music to be converted, extracts the second audio features of the reference music and the second sub-features of the music components based on the audio data of the reference music, constructs the sub-feature pairs belonging to different music components based on the music components of the respective first sub-features and the music components of the respective second sub-features, the sub-feature pair of any music component represents the feature mapping relationship of the second sub-feature belonging to the music component to the music to be converted, and generates the target music based on the sub-feature pairs, the first audio features and the second audio features. Therefore, the music of the desired style can be automatically converted based on the audio data of the music to be converted and the audio data of the reference music, the sub-feature pairs belonging to different music components are used as feature data for generating the target music, the feature matching relationship between the music to be converted and the reference music can be preserved as much as possible at the music component level, the granularity of the target music generation is improved, and the loss of detailed data in the conversion process can be reduced as much as possible for music with complex audio signals. Therefore, the automatic conversion of the music style can be realized, and the accuracy of the music conversion is improved.

[0065] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and the specific implementation can be referred to the description of the above method embodiments. For brevity, details are not repeated here.

[0066] The above description of each embodiment tends to emphasize the differences between the embodiments, and the same or similar parts can be mutually referred to for brevity. Details are not repeated here.

[0067] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus implementation described above is only schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual elements can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.

[0068] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.

[0069] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0070] When the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to perform all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various other media that can store program codes.

Claims

1. A music conversion method, characterized in that: include: Extracting, based on the audio data of the music to be converted, a first audio feature and first sub-features related to a plurality of musical components of the music to be converted, and extracting, based on the audio data of the reference music, a second audio feature and second sub-features related to a plurality of musical components of the reference music; Based on the music components of each of the first sub-features and the music components of each of the second sub-features, constructing sub-feature pairs belonging to different music components; wherein the sub-feature pair of any of the music components represents: a feature mapping relationship of the second sub-feature belonging to the music component with respect to the music to be converted; Target music is generated based on the sub-feature pairs, the first audio feature, and the second audio feature.

2. The method according to claim 1, characterized in that The generating target music based on the sub-feature pair, the first audio feature, and the second audio feature includes: Based on each of the sub-feature pairs, the first audio feature, and the second audio feature, respectively generate audio hidden layer features of the sub-feature pairs belonging to each of the music components; The target music is obtained by decoding based on the respective audio hidden layer features, the first audio features, and the second audio features.

3. The method according to claim 2, characterized in that Before respectively generating audio hidden layer features of the sub-feature pairs belonging to each of the music components based on the sub-feature pairs, the first audio features, and the second audio features, the method further includes: Based on the music component to which the sub-feature pair belongs, obtaining a first weight, a second weight, and a third weight; wherein the first weight is not less than the sum of the values ​​of the second weight and the third weight, the first weight represents the importance of the sub-feature pair in the audio hidden layer feature generation stage, the second weight represents the importance of the first audio feature in the audio hidden layer feature generation stage, and the third weight represents the importance of the second audio feature in the audio hidden layer feature generation stage; Generating audio hidden layer features of the sub-feature pairs belonging to each of the music components based on the sub-feature pairs, the first audio features, and the second audio features, respectively, includes: weighting the sub-feature pair, the first audio feature, and the second audio feature based on the first weight, the second weight, and the third weight corresponding to the music component to which the sub-feature pair belongs, respectively, to obtain weighted sub-feature pair, the first audio feature, and the second audio feature; Based on the weighted sub-feature pair, the first audio feature, and the second audio feature, audio hidden layer features of the sub-feature pair are obtained through processing.

4. The method according to claim 1, wherein The step of constructing sub-feature pairs belonging to different music components based on the music components of each of the first sub-features and the music components of each of the second sub-features includes: In response to the presence of the first sub-feature being consistent with the musical component of the second sub-feature, constructing the sub-feature pair based on the second sub-feature and the first sub-feature being consistent with the musical component of the second sub-feature; In response to the absence of the first sub-feature consistent with the musical component of the second sub-feature, the sub-feature pair is constructed based on the second sub-feature and a zero-vector feature.

5. The method according to claim 1, wherein The method further comprises: Obtaining target sub-audios of several musical components in target music; Based on the target sub-audio, encoding obtains target sub-features; Wherein, when the target music is the music to be converted, the target sub-feature is the first sub-feature; when the target music is the reference music, the target sub-feature is the second sub-feature.

6. The method according to claim 5, characterized in that The target sub-features are obtained by encoding the target music by a sub-component encoder, and the training process of the sub-component encoder includes: Acquire sample music containing a plurality of music component sample sub-audios, wherein each sample sub-audio is labeled with a sample music component; Acquire a positive sample audio and a negative sample audio about the sample music component; wherein the positive sample audio is consistent with the sample music component of the sample audio, and the negative sample audio is inconsistent with the sample music component of the sample audio; Encode the sample sub-audio, positive sample audio, and negative sample audio of the sample music component based on the sub-component encoder to obtain sample sub-audio features, positive sample audio features, and negative sample audio features; The differences among the sample sub-audio features, the positive sample audio features, and the negative sample audio features are measured based on a triplet loss function to adjust the network parameters of the sub-component encoder.

7. The method according to claim 1, characterized in that Before extracting the first audio feature of the music to be converted and first sub-features related to a plurality of music components based on the audio data of the music to be converted, the method further includes: Obtain a standard musical instrument code library; wherein the standard musical instrument code library contains standard codes of basic pitches of various musical instruments; In the case where the music component of the first sub-feature belongs to a musical instrument, after extracting a plurality of first sub-features related to the music to be converted and before constructing sub-feature pairs belonging to different music components, the method further includes: Selecting a first standard code in the standard instrument code library that is consistent with the first sub-feature music component; fusing the first sub-feature with the first standard code to obtain a fused code as the updated first sub-feature; In the case where the musical component of the second sub-feature belongs to a musical instrument, after extracting a plurality of second sub-features related to the reference music and before constructing sub-feature pairs belonging to different musical components, the method further includes: Selecting a second standard code in the standard instrument code library that is consistent with the second sub-feature music component; A fused code obtained by fusing the second sub-feature with the second standard code is used as the updated second sub-feature.

8. A music conversion device, characterized in that: include: an extraction module configured to extract, based on the audio data of the music to be converted, a first audio feature of the music to be converted and first sub-features related to a plurality of musical components, and, based on the audio data of the reference music, to extract a second audio feature of the reference music and second sub-features related to a plurality of musical components; a construction module, configured to construct sub-feature pairs belonging to different music components based on the music components of each of the first sub-features and the music components of each of the second sub-features; wherein the sub-feature pair of any of the music components represents: a feature mapping relationship of the second sub-feature belonging to the music component with respect to the music to be converted; A generating module is configured to generate target music based on the sub-feature pair, the first audio feature, and the second audio feature.

9. An electronic device, characterized in that: It comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the music conversion method according to any one of claims 1 to 7.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the music conversion method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Melody style conversion method and device, terminal equipment and storage medium

    CN113851098A

  • Singing conversion method and device, electronic equipment and storage medium

    CN119091854A