AI music generation method and system based on BPM rhythm matching

By extracting the MFCC, chromaticity and rhythm characteristics of music, combining user preferences and listening records, using the GAN model to generate AI music that meets the user's personalized and rhythm needs, solving the problems of incomplete music generation and unconsidered user preferences in the existing technology, and achieving high-quality AI music generation.

CN120375787AActive Publication Date: 2025-07-25XIAN HUACAI EXCELLENT TECHNOLOGY CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510503654.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-25
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing AI music generation methods fail to fully consider music characteristics and user personalized preferences, resulting in the generated music being unable to meet the unique needs of different users.

Method used

By extracting MFCC, chromaticity and rhythm characteristics, a two-dimensional classification matrix is constructed, combining user listening records and BPM requirements, a GAN model is used to generate AI music that meets user preferences and rhythm needs.

Benefits of technology

The generated AI music is more in line with the user's personalized needs, meets specific rhythm requirements, and provides high-quality music generation effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375787A_ABST
    Figure CN120375787A_ABST
Patent Text Reader

Abstract

The invention discloses an AI music generation method and system based on BPM rhythm matching, and relates to the technical field of artificial intelligence music generation, and the method comprises the steps: obtaining a user music library, constructing a two-dimensional classification matrix according to music types and emotion tone data, and summarizing the music according to combinations; framing the audio signal of each sub-music block, extracting MFCC, chroma and rhythm feature vectors, splicing and then combining a multi-frame mean value and a maximum value to obtain a sub-music block comprehensive feature vector; when an AI music generation request is received, sub-music blocks in a song listening record of a user are obtained, screening is carried out according to the BPM requirement, the weight is calculated according to the playing frequency, and a reference comprehensive feature vector is obtained; counting the comprehensive playing frequency of the secondary classification combination, and ranking and enqueueing; combinations are taken from the queue in sequence, sub-music blocks conforming to BPM are screened, a similarity index is calculated to screen candidate sub-music blocks, and AI music is generated based on a GAN model; the method for generating the AI music based on user preference and BPM rhythm matching is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence music generation, and particularly to an AI music generation method and system based on BPM rhythm matching. Background Art

[0002] With the rapid development of artificial intelligence technology, significant progress has been made in the field of AI music generation; traditional music creation mainly relies on the inspiration, expertise, and creative skills of human composers, and the creation process is time-consuming and limited by the creator's own experience and style; the emergence of AI music generation technology has brought new possibilities to music creation; it uses algorithms such as machine learning and deep learning, and through the learning and analysis of a large amount of music data, it can automatically generate music works with a certain style and expressiveness;

[0003] Currently, AI music generation technology has been applied in multiple fields, such as game music creation, film and television scoring, advertising music production, etc.; however, existing AI music generation methods still have some deficiencies. For example: existing methods may not be comprehensive enough in music feature extraction, only considering some features, such as rhythm or pitch, while ignoring other important music features, which may limit the understanding and expression ability of AI music generation systems for music content; secondly, existing methods may not fully consider the personalized preferences of users, such as music category, emotional tone, and listening habits, which makes it difficult for AI music generation systems to meet the unique preferences of different users. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] In view of the technical problems in the background art, the present invention proposes an AI music generation method and system based on BPM rhythm matching. By extracting MFCC features, chroma features, and rhythm features, and performing multi-dimensional feature fusion, candidate sub-music blocks are screened through the calculation of weight coefficients and similarity indices, in combination with the user's music category preferences, emotional tone preferences, and listening record data; thus solving the technical problems recorded in the background art.

[0006] (2) Technical Solutions

[0007] To achieve the above object, the present invention is realized through the following technical solutions:

[0008] An AI music generation method based on BPM rhythm matching, comprising:

[0009] Obtain the music library provided by the user, construct a two-dimensional classification matrix based on music category data and emotional tone data, and summarize the music in the music library according to the two-dimensional classification combination; perform chunking operations on each piece of music, and mark the BPM value and the category of the two-dimensional classification combination for the generated several sub-music chunks;

[0010] Perform frame segmentation on the audio signal corresponding to each sub-music chunk respectively, and extract MFCC features, chroma features, and rhythm features to obtain the MFCC feature vector, chroma feature vector, and rhythm feature vector for each frame; splice the three feature vectors to obtain the high-dimensional feature vector for each frame; combine the mean and maximum values of the feature vectors of multiple frames to obtain the comprehensive feature vector for each sub-music chunk;

[0011] When an AI music generation request is received, obtain the sub-music chunks of each music in the user's listening record data, filter the sub-music chunks based on the required BPM value, and calculate the weight coefficient of each corresponding sub-music chunk based on the play count of each music. After combination, obtain the reference comprehensive feature vector under the corresponding required BPM value; based on the user's listening record data, count and calculate the comprehensive play frequency of each secondary classification combination, and after sorting, put it into the matching waiting queue;

[0012] Successively obtain a secondary classification combination from the matching waiting queue, filter out all sub-music chunks that meet the required BPM value under this secondary classification combination, successively obtain the comprehensive feature vectors of the sub-music, calculate the similarity index with the reference comprehensive feature vector respectively, perform screening of candidate sub-music chunks, and generate AI music based on the GAN model and in combination with all candidate sub-music chunks.

[0013] Specifically, there are n types of music category data and m types of emotional tone data in the music library. Combine all the music category data and emotional tone data in the music library to construct a two-dimensional classification matrix, and obtain n×m groups of secondary classification combinations; summarize all the music in all the music libraries provided by users, and match each music with the corresponding secondary classification combination.

[0014] Specifically, increase the energy of the high-frequency part through a high-pass filter to perform pre-emphasis operation; perform frame segmentation on the audio signal corresponding to each sub-music chunk; apply a window function to each frame of the segmented signal for windowing operation;

[0015] Perform a fast Fourier transform on the single-frame audio signal after pre-emphasis and windowing to obtain the frequency-domain signal of each frame of the audio signal;

[0016] Calculate the power spectrum of each frame to quantify the spectral energy distribution;

[0017] Map the spectrum to the Mel frequency scale, and perform weighted averaging on the spectrum through a set of triangular filters, and take the logarithm of each value output by the filter bank;

[0018] Perform a discrete cosine transform on the logarithmic Mel spectrogram to obtain Mel frequency cepstral coefficients, and obtain a 13-dimensional MFCC feature vector.

[0019] Specifically, perform frame segmentation on the audio signal corresponding to each sub-music block, and calculate the short-time energy of each frame;

[0020] Based on the difference of the energy of two adjacent frames, determine whether each frame is a starting point, and determine the starting point intensity;

[0021] Analyze the beat position of the audio signal based on the autocorrelation method, and calculate the beat frequency of each frame;

[0022] Combine the energy feature, starting point intensity, and beat probability of each frame in the audio signal to form a 3-dimensional rhythm feature vector.

[0023] Furthermore, perform frame segmentation on the audio signal corresponding to each sub-music block, calculate the short-time Fourier transform of the audio signal to obtain the spectrum, divide the spectrum into 12 sub-bands, and calculate the energy of each sub-band, which is the 12-dimensional chroma feature vector;

[0024] Concatenate the 13-dimensional MFCC feature vector, 12-dimensional chroma feature vector, and 3-dimensional rhythm feature vector of each frame in the audio signal to obtain a 28-dimensional feature vector for each frame;

[0025] Perform a mean operation on the corresponding positions of the 28-dimensional feature vectors of all frames in the audio signal to obtain a 28-dimensional first feature vector; obtain the maximum value at each position in the 28-dimensional feature vectors of all frames in the audio signal, and combine them in order to obtain a 28-dimensional second feature vector; concatenate the 28-dimensional first feature vector and the 28-dimensional second feature vector to obtain a 56-dimensional comprehensive feature vector corresponding to the audio signal, that is, the comprehensive feature vector corresponding to each sub-music block.

[0026] Specifically, obtain the music corresponding to the listening record data, as well as the music category and emotional tone data of each music; respectively obtain several sub-music blocks obtained after dividing each music;

[0027] Obtain the BPM requirement value in the AI music generation request, screen out the sub-music blocks with the same BPM value as the BPM requirement value from the obtained several sub-music blocks, and perform the weight coefficient λ i calculation, specifically: where i represents the i-th sub-music block screened out, c i represents the number of plays of the music to which the i-th sub-music block belongs, and M represents the total number of sub-music blocks screened out;

[0028] Obtain the comprehensive feature vectors of all the screened-out sub-music blocks and based on the weight λ of each sub - music block i , calculate the reference comprehensive feature vector of the current BPM requirement value The expression is:

[0029] Furthermore, based on the music category and emotional tone data of each music in the listening record data, respectively count the music playing frequency corresponding to each music category, and the music playing frequency corresponding to each emotional tone;

[0030] Add the music playing frequency corresponding to any music category to the music playing frequency corresponding to any emotional tone to obtain the comprehensive playing frequency of the corresponding secondary classification combination. Similarly, calculate the comprehensive playing frequency of each secondary classification combination in turn, and sort all the secondary classification combinations in descending order of the comprehensive playing frequency, and put them into the matching waiting queue in order.

[0031] Furthermore, sequentially obtain a secondary classification combination from the matching waiting queue, and obtain all the sub - music blocks under this secondary classification combination;

[0032] Screen out the sub - music blocks with the BPM value equal to the BPM requirement value from all the obtained sub - music blocks, and calculate the comprehensive feature vector of each screened sub - music block based on the Euclidean distance and the reference comprehensive feature vector to get the similarity index Si j ;

[0033] Sort the similarity index Si of the comprehensive feature vector of each screened sub - music block and the reference comprehensive feature vector j in descending order; set the similarity index threshold Si0, and sequentially select several sub - music blocks with the similarity index Si j greater than the similarity index threshold Si0 as candidate sub - music blocks.

[0034] Specifically, denote the form of the comprehensive feature vector as {x j1 , x j2 , … x j56}, and the form of the reference comprehensive feature vector as {x1, x2, … x 56};

[0035] The expression of the similarity index Si j is: Si j =α1 * MSi j +α2 * CSi j+α3*RSi j , where α1, α2, and α3 represent the weight coefficients of the MFCC feature vector, the chroma feature vector, and the rhythm feature vector respectively, and α1 + α2 + α3 = 1;

[0036] MSi j is the similarity index of the MFCC feature vector, and the expression is

[0037] CSi j is the similarity index of the chroma feature vector, and the expression is:

[0038] RSi j is the similarity index of the rhythm feature vector, and the expression is:

[0039] An AI music generation system based on BPM rhythm matching, comprising:

[0040] A music classification and summarization module, which is used to obtain the music library provided by the user, construct a two-dimensional classification matrix based on music category data and emotional tone data, and summarize the music in the music library according to the two-dimensional classification combination; perform a chunking operation on each piece of music, and mark the BPM value and the category of the two-dimensional classification combination for the generated several sub-music chunks;

[0041] A feature extraction and processing module, which is used to frame the audio signal corresponding to each sub-music chunk respectively, and extract the MFCC feature, chroma feature, and rhythm feature frame by frame to obtain the MFCC feature vector, chroma feature vector, and rhythm feature vector of each frame; splice the three feature vectors to obtain the high-dimensional feature vector of each frame; combine the mean and maximum values of the feature vectors of multiple frames to obtain the comprehensive feature vector of each sub-music chunk;

[0042] An AI music matching module, when receiving an AI music generation request, is used to obtain the sub-music chunks of each music in the user's listening record data, screen the sub-music chunks based on the BPM required value, and calculate the weight coefficient of each corresponding sub-music chunk based on the playback times of each music, and combine them to obtain the reference comprehensive feature vector under the corresponding BPM required value; based on the user's listening record data, count and calculate the comprehensive playback frequency of each secondary classification combination, and put them into the matching waiting queue after sorting;

[0043] An AI music generation module, which is used to sequentially obtain a secondary classification combination from the matching waiting queue, screen out all sub-music chunks that meet the BPM required value under this secondary classification combination, sequentially obtain the comprehensive feature vectors of the sub-music, calculate the similarity index with the reference comprehensive feature vector respectively, screen the candidate sub-music chunks, and generate AI music based on the GAN model and in combination with all the candidate sub-music chunks.

[0044] (III) Beneficial effects

[0045] The present invention provides an AI music generation method and system based on BPM rhythm matching, which has the following beneficial effects:

[0046] 1. By obtaining the user's music library, constructing a two-dimensional classification matrix based on music categories and emotional tone data, summarizing the music, and then dividing the music into blocks and marking the BPM values and classification combination categories; it can accurately integrate the user's music preferences, provide a rich and orderly music material library for subsequent generation of AI music based on user preferences and rhythm matching, and make the generated AI music more in line with the user's personalized needs;

[0047] 2. By dividing each sub-music block into frames, extracting MFCC, chroma, and rhythm features frame by frame, splicing them to obtain a high-dimensional feature vector, and then combining the multi-frame mean and maximum value to obtain a comprehensive feature vector; it comprehensively and meticulously mines the feature information of the music signal, makes the generated AI music more in line with the target requirements at the feature level, provides an accurate and effective feature basis for subsequent music similarity calculation and AI music generation, and helps to generate high-quality AI music that matches the reference features;

[0048] 3. After receiving an AI music generation request, based on the user's listening record and the required BPM value, screening sub-music blocks, calculating the weight coefficient to obtain a reference comprehensive feature vector, statistically sorting the secondary classification combination playback frequencies and putting them into a matching queue, screening candidate sub-music blocks and generating AI music based on the GAN model; comprehensively considering the user's recent listening habits and rhythm requirements, generating AI music by reasonably screening and matching music blocks and using the GAN model, so that the generated AI music not only meets the user's preferences but also satisfies specific rhythm requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a step schematic diagram of an AI music generation method based on BPM rhythm matching provided by the present invention;

[0050] Figure 2 It is a structural schematic diagram of an AI music generation system based on BPM rhythm matching provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] Reference Figure 1, the present invention provides an AI music generation method based on BPM rhythm matching, including:

[0053] Step 1: Obtain the music library provided by the user, construct a two-dimensional classification matrix based on music category data and emotional tone data, and summarize the music in the music library according to the two-dimensional classification combination; perform a chunking operation on each piece of music, and mark the BPM value and the category of the two-dimensional classification combination for the generated several sub-music chunks;

[0054] The first step includes the following steps:

[0055] Step 101: Number each user with an id, and obtain the music library provided by each user; the music library stores all the music that the user hopes to fuse. The present invention needs to compare the music libraries provided by several users in combination with the user's preferences, and then create AI music that meets the user's preferences;

[0056] The music library stores the music category data and emotional tone data of each piece of music. The music category data refers to the information obtained by classifying and labeling music works according to styles, genres or types, such as rock, pop, classical, jazz, etc., and is used to describe the overall style and characteristics of the music; there are n types of music category data, and specifically use A a to represent the a-th type of music category data;

[0057] The emotional tone data refers to the information describing the emotional tendency or atmosphere expressed by the music work, such as cheerful, sad, exciting, calm, etc., and is used to convey the emotional experience brought by the music to the listener; there are m types of emotional tone data, and specifically use B b to represent the b-th type of emotional tone data;

[0058] Each piece of music belongs to one type of music category data and one type of emotional tone data at the same time;

[0059] Step 102: Summarize all the music libraries provided by all users, and perform a secondary classification summary on all the music according to different music categories and emotional tone combinations, specifically:

[0060] Combine all the music category data and emotional tone data and construct a two-dimensional classification matrix to obtain n×m groups of secondary classification combinations;

[0061] Summarize all the music in all the music libraries provided by all users, match each piece of music with the corresponding secondary classification combination, and mark each piece of music with the same mark as the corresponding user who provided it. The same piece of music can have marks of multiple users, that is, the same piece of music exists in the music libraries provided by multiple users at the same time;

[0062] Step 103: Sequentially obtain each piece of music under each secondary classification combination, and divide each obtained piece of music into chunks, specifically into 30 - second sub - music chunks. If the duration of the last sub - music chunk after dividing a piece of music is less than 30 seconds, extend this sub - music chunk. Specifically, extend 30 seconds forward from the end of this piece of music, and use the 30 - second corresponding sub - music chunk as the last sub - music chunk of the corresponding music;

[0063] Mark each sub - music chunk with its corresponding music name, and use audio analysis software to obtain the BPM value of each sub - music chunk; according to the music of each sub - music chunk, make each sub - music chunk correspond one - by - one with the secondary classification combination.

[0064] When in use, combine the content in Steps 101 to 103:

[0065] By obtaining the user's music library, constructing a two - dimensional classification matrix based on music categories and emotional tone data and summarizing the music, then dividing the music into chunks and marking the BPM value and classification combination category; it can accurately integrate the user's music preferences, provide a rich and orderly music material library for subsequent generation of AI music based on user preferences and rhythm matching, and make the generated AI music more in line with the user's personalized needs.

[0066] Step Two: Respectively frame the audio signals corresponding to each sub - music chunk, and extract MFCC features, chroma features, and rhythm features frame by frame to obtain the MFCC feature vector, chroma feature vector, and rhythm feature vector of each frame; splice the three feature vectors to obtain the high - dimensional feature vector of each frame; combine the mean and maximum value of the feature vectors of multiple frames to obtain the comprehensive feature vector of each sub - music chunk.

[0067] The steps in Step Two include the following steps:

[0068] Step 201: Extract music features from the audio signals corresponding to each sub - music chunk. The music features specifically include Mel - Frequency Cepstral Coefficient features, that is, MFCC features, chroma features, and rhythm features;

[0069] Step 202: MFCC features are used to capture the spectral features of audio signals. Based on the non - linear perception of frequency by the human ear (Mel scale), extract the features of audio signals. The specific steps are as follows:

[0070] Step 2021: Increase the energy of the high - frequency part through a high - pass filter to achieve pre - emphasis operation. This can enhance the high - frequency components in the audio signal, compensate for the possible attenuation of high - frequency signals during transmission, and make the spectrum flatter. The expression is: s l '(k) = s l (k)-β·s l (k - 1), where sl '(k) represents the pre-emphasized signal, s l (k) represents the k-th sampling point of the l-th frame, and β represents the pre-emphasis coefficient;

[0071] Step 2022: Perform frame segmentation on the audio signal corresponding to each sub-music block. Specifically, divide the audio signal into several short-time frames with a frame length of 30 ms, and make there be a 50% overlap between adjacent frames;

[0072] Step 2023: Apply a window function to each frame of the signal for windowing operation to reduce spectral leakage and make the frequency-domain analysis more accurate; the window function is specifically the Hamming window function, and the expression is: where N represents the window length, that is, the number of sampling points, and w(k) represents the window function value; multiply the window function by each frame of the signal point by point, and the expression is:

[0073] Step 2024: Perform a fast Fourier transform on the single-frame audio signal after pre-emphasis and windowing to obtain the frequency-domain signal of each frame of the audio signal, and the expression is: where S l (o) represents the frequency number of the l-th frame, and I represents the imaginary unit;

[0074] Step 2025: Calculate the power spectrum of each frame to quantify the spectral energy distribution, and the expression is: where P l (o) represents the power spectrum of the l-th frame;

[0075] Step 2026: Map the spectrum to the Mel frequency scale and perform weighted averaging on the spectrum through a set of triangular filters; the relationship between the Mel frequency and the ordinary frequency f is:

[0076] Take the logarithm of each value output by the filter bank to simulate the human ear's perception of the sound intensity;

[0077] Perform a discrete cosine transform on the logarithmic Mel spectrum to obtain the Mel frequency cepstral coefficients. Since the low-order coefficients contain the main spectral information and the high-order coefficients are often related to noise, usually the first 12 to 13 coefficients are retained. Here, the first 13 coefficients are taken, that is, a 13-dimensional MFCC feature vector is obtained;

[0078] Step 203: The chroma feature reflects the color characteristics of the audio signal that do not contain brightness information and is used to describe the tone and saturation of the audio signal;

[0079] Specifically, the audio signal is also segmented into several short-time frames with a frame length of 30 ms. The spectrum is obtained by calculating the short-time Fourier transform of the audio signal, and then the spectrum is divided into 12 sub-bands. The energy of each sub-band is calculated and normalized to eliminate the influence of volume changes. Finally, the obtained 12-dimensional chromatic feature vector represents the energy distribution of the audio signal on different sub-bands.

[0080] Step 204: The rhythm feature is a time-related feature in the audio signal, specifically including energy feature, onset intensity, and beat probability. It describes information such as the beat intensity and rhythm change of the audio signal. Specifically:

[0081] Step 2041: Specifically, the audio signal is also segmented into several short-time frames with a frame length of 30 ms, and the short-time energy E(l) of each frame is calculated. The expression is: where s(l + k) represents the value of the audio signal at the (l + k)-th sampling point;

[0082] Step 2042: Detect whether each frame is an onset point (a mutation point of energy or spectrum). First, calculate the difference D(l) between the energies of two adjacent frames. The expression is: D(l) = E(l) - E(l - 1); Set a difference threshold. If the difference D(l) between the energies of two adjacent frames exceeds the difference threshold, then mark the l-th frame as an onset point; For each frame, if it is an onset point, then the onset intensity S(l) of this frame = D(l); Otherwise, the onset intensity S(l) of this frame = 0.

[0083] Step 2043: Estimate the beat position of the audio signal based on the autocorrelation method, and calculate the autocorrelation function R(τ). The expression is: where τ represents the delay amount, that is, the time offset when the signal is compared with itself; Calculate R(τ) sequentially starting from τ = 1 to obtain the autocorrelation sequence; Find the first peak in the autocorrelation sequence, and use the delay amount corresponding to this peak as the beat period T, that is, estimate the beat position of the audio signal as K·T, and K ∈ {1, 2,...};

[0084] Calculate the beat probability P(l) of each frame. The expression is where l is the index of the current frame, l near is the index of the nearest beat to the current frame, and σ represents the standard deviation, and the specific value is one-fourth of the beat interval;

[0085] Step 2044: Combine the energy feature, onset intensity, and beat probability of each frame in the audio signal to form a 3-dimensional rhythm feature vector.

[0086] Step 205: Concatenate the 13-dimensional MFCC feature vectors, 12-dimensional chroma feature vectors, and 3-dimensional rhythm feature vectors of each frame in the audio signal to obtain a 28-dimensional feature vector for each frame;

[0087] Perform a mean operation on the corresponding positions of the 28-dimensional feature vectors of all frames in the audio signal to obtain a 28-dimensional first feature vector; Obtain the maximum value at each position in the 28-dimensional feature vectors of all frames in the audio signal, and combine them in order to obtain a 28-dimensional second feature vector; Concatenate the 28-dimensional first feature vector and the 28-dimensional second feature vector to obtain a 56-dimensional comprehensive feature vector corresponding to the audio signal;

[0088] Analyze and obtain the comprehensive feature vectors of each sub-music block based on the above steps.

[0089] When in use, combine the content in Steps 201 to 205:

[0090] By performing frame segmentation on each sub-music block, extracting MFCC, chroma, and rhythm features frame by frame, concatenating them to obtain high-dimensional feature vectors, and then combining the multi-frame mean and maximum values to obtain comprehensive feature vectors; Comprehensively and meticulously excavate the feature information of the music signal, making the generated AI music more in line with the target requirements at the feature level, providing accurate and effective feature basis for subsequent music similarity calculation and AI music generation, and helping to generate high-quality AI music that matches the reference features.

[0091] Step Three: When receiving an AI music generation request, obtain the sub-music blocks of each music in the user's music listening record data, screen the sub-music blocks based on the BPM requirement value, and calculate the weight coefficient of each corresponding sub-music block based on the play count of each music. After combining, obtain the reference comprehensive feature vector under the corresponding BPM requirement value; Based on the user's music listening record data, count and calculate the comprehensive play frequency of each secondary classification combination, sort them and put them into the matching waiting queue; Sequentially obtain a secondary classification combination from the matching waiting queue, screen out all sub-music blocks that meet the BPM requirement value under this secondary classification combination, sequentially obtain the comprehensive feature vectors of the sub-music and calculate the similarity index with the reference comprehensive feature vector respectively, perform screening of candidate sub-music blocks, and generate AI music based on the GAN model and in combination with all candidate sub-music blocks;

[0092] The steps in Step Three include the following steps:

[0093] Step 301: When the user sends an AI music generation request, obtain the music listening record data of this user. The music corresponding to the music listening record data of each user comes from the music library provided by the user;

[0094] Obtain all the music played in the recent several days and the corresponding play count data from the music listening record data of this user;

[0095] Step 302: Obtain the music corresponding to the music listening record data, as well as the music category and emotional tone data of each music; respectively obtain a number of sub-music blocks obtained after each music is segmented;

[0096] Obtain the BPM required value in the AI music generation request, screen out the sub-music blocks with the same BPM value as the BPM required value from the obtained number of sub-music blocks, and perform the calculation of the weight coefficient λ i Specifically: where i represents the i-th sub-music block screened out, c i represents the number of plays of the music to which the i-th sub-music block belongs, and M represents the total number of sub-music blocks screened out; the weight coefficient represents the reference index of each sub-music block. When generating AI music, the user's recent music listening habits need to be considered to intelligently generate AI music that the user is most likely to be interested in;

[0097] Step 303: Obtain the comprehensive feature vector of all the screened-out sub-music blocks and calculate the reference comprehensive feature vector of the current BPM required value based on the weight λ i of each sub-music block The expression is:

[0098] Step 304: Based on the music category and emotional tone data of each music in the music listening record data, respectively count the music play frequencies corresponding to each music category, and the music play frequencies corresponding to each emotional tone. Specifically, take the sum of the number of plays of all the music under each music category in the music listening record data as the music play frequency corresponding to the corresponding music category, and the statistical method of the music play frequency corresponding to each emotional tone is the same;

[0099] Add the music play frequency corresponding to any music category to the music play frequency corresponding to any emotional tone to obtain the comprehensive play frequency of the corresponding secondary classification combination. Similarly, calculate the comprehensive play frequencies of each secondary classification combination in turn, and sort all the secondary classification combinations in descending order of the comprehensive play frequency, and put them into the matching waiting queue in order;

[0100] Step 305: Obtain a secondary classification combination from the matching waiting queue in turn, and obtain all the sub-music blocks under this secondary classification combination. The sub-music blocks obtained here include a number of sub-music blocks in all the music libraries provided by the user. Combining the music libraries of multiple users helps to generate more reasonable AI music;

[0101] Filter out the sub-music blocks with the same BPM value as the required BPM value from all the obtained sub-music blocks, and calculate the comprehensive feature vector of each filtered sub-music block based on the Euclidean distance with the reference comprehensive feature vector for the similarity index Si j , and denote the comprehensive feature vector in the form of {x j1 , x j2 , … x j56}, and the reference comprehensive feature vector in the form of {x1, x2, … x 56}, where the elements of the 1st to 13th items, and the 29th to 41st items are the elements of the MFCC feature vector, then the similarity index MSi j of the MFCC feature vector is

[0102] where the elements of the 14th to 25th items, and the 42nd to 53rd items are the elements of the chroma feature vector, then the similarity index CSi j of the chroma feature vector is;

[0103] where the elements of the 26th to 28th items, and the 54th to 56th items are the elements of the rhythm feature vector, then the similarity index RSi j of the rhythm feature vector is;

[0104] Combining the similarity index MSi j of the MFCC feature vector, the similarity index CSi j of the chroma feature vector, and the similarity index RSi j of the rhythm feature vector, the similarity index Si of the comprehensive feature vector and the reference comprehensive feature vector is obtained, and the expression is: Si j = α1 * MSi j + α2 * CSi j + α3 * RSi j + α3 * RSi j , where α1, α2, and α3 respectively represent the weight coefficients of the MFCC feature vector, the chroma feature vector, and the rhythm feature vector, and the specific values are determined by the user's own input, and α1 + α2 + α3 = 1;

[0105] Step 306. The similarity index Si of the comprehensive feature vector of each filtered sub-music block and the reference comprehensive feature vector jSort them in descending order; set the similarity index threshold Si0, and sequentially select the similarity index Si in the order of secondary classification combinations j Select several sub-music blocks with similarity indices greater than the similarity index threshold Si0 as candidate sub-music blocks; the number of sub-music blocks selected is determined by the user himself, and the length of one sub-music block is 30 seconds; if the candidate sub-music blocks selected in the current secondary classification combination do not reach the number of sub-music blocks selected by the user, then sequentially select the next secondary classification combination in the matching waiting queue for the selection of candidate sub-music blocks;

[0106] Take the comprehensive feature vector corresponding to each candidate sub-music block as the input condition of the GAN model; construct a generator to receive a random noise vector and output a generated music segment, construct a discriminator to receive the music segment and output the probability value that the segment is real; randomly initialize the parameters of the generator and the discriminator, and perform adversarial training, specifically: the generator generates a music segment and passes it to the discriminator, the discriminator judges the authenticity of the music segment and calculates the loss, and updates the parameters of the generator and the discriminator according to the loss; repeat the above adversarial training process until the model converges; take the matched music features as the input condition of the generator, and the generator generates a new music segment according to the input features; splice the music segments generated by each candidate sub-music block to obtain the intelligently generated AI music.

[0107] When in use, combine the content in steps 301 to 306:

[0108] After receiving an AI music generation request, screen sub-music blocks based on the user's listening record and the BPM requirement value, calculate the weight coefficient to obtain a reference comprehensive feature vector, count the playback frequency of the secondary classification combinations and sort them into the matching queue, screen candidate sub-music blocks and generate AI music based on the GAN model; comprehensively consider the user's recent listening habits and rhythm requirements, reasonably screen and match music blocks, and use the GAN model to generate AI music, so that the generated AI music meets both the user's preferences and specific rhythm requirements.

[0109] Reference Figure 2 , the present invention also provides an AI music generation system based on BPM rhythm matching, including:

[0110] A music classification and summary module, used to obtain the music library provided by the user, construct a two-dimensional classification matrix based on music category data and emotional tone data, and summarize the music in the music library according to two-dimensional classification combinations; perform a block operation on each piece of music, and mark the BPM value and the category of the two-dimensional classification combination for the generated several sub-music blocks;

[0111] A feature extraction processing module is used to frame the audio signals corresponding to each sub - music block respectively, and extract MFCC features, chroma features and rhythm features frame by frame to obtain the MFCC feature vectors, chroma feature vectors and rhythm feature vectors of each frame; splice the three feature vectors to obtain the high - dimensional feature vectors of each frame; combine the mean and maximum values of the feature vectors of multiple frames to obtain the comprehensive feature vectors of each sub - music block.

[0112] An AI music matching module, when receiving an AI music generation request, is used to obtain the sub - music blocks of each music in the user's music listening record data, screen the sub - music blocks based on the required BPM value and calculate the weight coefficient of each corresponding sub - music block based on the play times of each music, and combine them to obtain the reference comprehensive feature vector under the required BPM value; based on the user's music listening record data, count and calculate the comprehensive play frequency of each secondary classification combination, and after sorting, put it into the matching waiting queue.

[0113] An AI music generation module is used to sequentially obtain a secondary classification combination from the matching waiting queue, screen out all sub - music blocks that meet the required BPM value under this secondary classification combination, sequentially obtain the comprehensive feature vectors of the sub - music and calculate the similarity index with the reference comprehensive feature vector respectively to screen the candidate sub - music blocks, and generate AI music based on the GAN model and in combination with all the candidate sub - music blocks.

[0114] In the above - mentioned embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general - purpose computer, a special - purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer storage medium or transmitted through a computer storage medium.

[0115] The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center in a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) manner. The computer storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. that contains one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid - state drive (SSD)), etc.

[0116] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.

Claims

1. An AI music generation method based on BPM rhythm matching, characterized in that: It includes the following steps: Obtain the music library provided by the user, construct a two-dimensional classification matrix based on music category data and emotional tone data, and summarize the music in the music library according to the two-dimensional classification combination; perform a chunking operation on each piece of music, and mark the BPM value and the category of the two-dimensional classification combination for the generated several sub-music chunks; Frame the audio signal corresponding to each sub-music chunk respectively, and extract MFCC features, chroma features, and rhythm features to obtain the MFCC feature vector, chroma feature vector, and rhythm feature vector for each frame; splice the three feature vectors to obtain the high-dimensional feature vector for each frame; combine the mean and maximum value of the feature vectors of multiple frames to obtain the comprehensive feature vector for each sub-music chunk; When receiving an AI music generation request, obtain the sub-music chunks of each music in the user's listening record data, filter the sub-music chunks based on the required BPM value, calculate the weight coefficient for each corresponding sub-music chunk based on the play count of each music, and combine them to obtain the reference comprehensive feature vector under the required BPM value; based on the user's listening record data, count and calculate the comprehensive play frequency of each secondary classification combination, and put them into the matching waiting queue after sorting; Successively obtain a secondary classification combination from the matching waiting queue, filter out all sub-music chunks that meet the required BPM value under this secondary classification combination, successively obtain the comprehensive feature vectors of the sub-music, calculate the similarity index with the reference comprehensive feature vector respectively, perform screening of candidate sub-music chunks, and generate AI music based on the GAN model and in combination with all candidate sub-music chunks.

2. An AI music generation method based on BPM rhythm matching as described in claim 1, characterized in that: There are n kinds of music category data and m kinds of emotional tone data in the music library. Combine all the music category data and emotional tone data in the music library to construct a two-dimensional classification matrix, obtaining n×m groups of secondary classification combinations; summarize all the music in the music library provided by all users, and match each music with the corresponding secondary classification combination.

3. An AI music generation method based on BPM rhythm matching as described in claim 1, characterized in that: Increase the energy of the high-frequency part through a high-pass filter and perform pre-emphasis operation; perform a framing operation on the audio signal corresponding to each sub-music chunk; apply a window function to each framed signal for windowing operation; Perform a fast Fourier transform on the pre-emphasized and windowed single-frame audio signal to obtain the frequency-domain signal of each frame of the audio signal; Calculate the power spectrum of each frame to quantify the spectral energy distribution; Map the spectrum to the Mel frequency scale, perform weighted averaging on the spectrum through a group of triangular filters, and take the logarithm of each value output by the filter bank; Perform a discrete cosine transform on the logarithmic Mel spectrum to obtain the Mel frequency cepstral coefficients, and obtain a 13-dimensional MFCC feature vector.

4. An AI music generation method based on BPM rhythm matching as described in claim 3, characterized in that: Perform a framing operation on the audio signal corresponding to each sub-music chunk, and calculate the short-time energy of each frame; Based on the difference of energy between adjacent two frames, determine whether each frame is a starting point and determine the starting point intensity; Analyze the beat position of the audio signal based on the autocorrelation method and calculate the beat frequency of each frame; Combine the energy feature, starting point intensity, and beat probability of each frame in the audio signal to form a 3D rhythm feature vector.

5. A method for generating AI music based on BPM rhythm matching as described in claim 4, characterized in that: Perform frame division on the audio signal corresponding to each sub-music block, calculate the short-time Fourier transform of the audio signal to obtain the spectrum, divide the spectrum into 12 sub-bands, and calculate the energy of each sub-band, which is the 12D chroma feature vector; Concatenate the 13D MFCC feature vector, 12D chroma feature vector, and 3D rhythm feature vector of each frame in the audio signal to obtain a 28D feature vector for each frame; Perform a mean operation on the corresponding positions of the 28D feature vectors of all frames in the audio signal to obtain a 28D first feature vector; obtain the maximum value at each position in the 28D feature vectors of all frames in the audio signal, and combine them in order to obtain a 28D second feature vector; concatenate the 28D first feature vector and the 28D second feature vector to obtain a 56D comprehensive feature vector corresponding to the audio signal, that is, the comprehensive feature vector corresponding to each sub-music block.

6. A method for generating AI music based on BPM rhythm matching as described in claim 5, characterized in that: Obtain the music corresponding to the listening record data, as well as the music category and emotional tone data of each music; respectively obtain several sub-music blocks obtained after dividing each music; Obtain the BPM requirement value in the AI music generation request, screen out the sub-music blocks with the same BPM value as the BPM requirement value from the obtained several sub-music blocks, and perform the calculation of the weight coefficient λ i for the screened sub-music blocks, specifically: where i represents the i-th sub-music block screened out, and c i represents the number of playbacks of the music to which the i-th sub-music block belongs, and M represents the total number of sub-music blocks screened out; Obtain the comprehensive feature vectors of all the filtered sub-music blocks And based on the weight λ of each sub-music block i , calculate the reference comprehensive feature vector of the current BPM requirement value The expression is as follows:

7. A method for generating AI music based on BPM rhythm matching as described in claim 6, characterized in that: Based on the music category and emotional tone data of each music in the listening record data, respectively count the music playback frequencies corresponding to each music category, and the music playback frequencies corresponding to each emotional tone; Add the music playback frequency corresponding to any music category to the music playback frequency corresponding to any emotional tone to obtain the comprehensive playback frequency corresponding to the corresponding secondary classification combination. Similarly, calculate the comprehensive playback frequencies of each secondary classification combination in turn, sort all the secondary classification combinations in descending order of the comprehensive playback frequency, and place them in the matching waiting queue in order.

8. A method for generating AI music based on BPM rhythm matching as described in claim 7, characterized in that: Successively obtain a secondary classification combination from the matching waiting queue, and obtain all the sub-music blocks under this secondary classification combination; Select sub-music chunks with BPM values equal to the required BPM value from all the obtained sub-music chunks, and calculate the comprehensive feature vector of each selected sub-music chunk based on the Euclidean distance and the reference comprehensive feature vector for the similarity index Si j ; The comprehensive feature vectors of each filtered sub-music block and the reference comprehensive feature vector for the similarity index Si j are sorted in descending order; a similarity index threshold Si0 is set, and the similarity index Si j for several sub-music blocks greater than the similarity index threshold Si0 are selected in the order of secondary classification combinations as candidate sub-music blocks.

9. A method for generating AI music based on BPM rhythm matching as described in claim 8, characterized in that: Denote the comprehensive feature vector in the form of {x j1 , x j2 , … x j56}, and the reference comprehensive feature vector in the form of {x1, x2, … x 56}; The similarity index Si j is expressed as: Si j = α1 * MSi j + α2 * CSi j + α3 * RSi j , where α1, α2, and α3 respectively represent the weight coefficients of the MFCC feature vector, the chroma feature vector, and the rhythm feature vector, and α1 + α2 + α3 = 1; MSi j is the similarity index of the MFCC feature vector, and the expression is CSi j Is the similarity index of the chromaticity feature vector, and the expression is: RSi j Is the similarity index of the rhythm feature vector, and the expression is:

10. An AI music generation system based on BPM rhythm matching, characterized in that, Includes: A music classification summary module, used to obtain the music library provided by the user, construct a two-dimensional classification matrix based on the music category data and emotional tone data, and summarize the music in the music library according to the two-dimensional classification combination; perform block division on each piece of music, and mark the BPM value and two-dimensional classification combination category of the generated several sub-music blocks; A feature extraction processing module, which is used to frame the audio signals corresponding to each sub-music block respectively, and extract MFCC features, chroma features and rhythm features frame by frame to obtain the MFCC feature vectors, chroma feature vectors and rhythm feature vectors of each frame; splice the three feature vectors to obtain the high-dimensional feature vectors of each frame; combine the mean and maximum values of the feature vectors of multiple frames to obtain the comprehensive feature vectors of each sub-music block; An AI music matching module, when receiving an AI music generation request, is used to obtain the sub-music blocks of each music in the user's music listening record data, filter the sub-music blocks based on the required BPM value, calculate the weight coefficient of each corresponding sub-music block based on the playback times of each music, and combine them to obtain the reference comprehensive feature vector under the required BPM value; based on the user's music listening record data, count and calculate the comprehensive playback frequency of each secondary classification combination, and put them into the matching waiting queue after sorting; An AI music generation module, which is used to sequentially obtain a secondary classification combination from the matching waiting queue, filter out all sub-music blocks that meet the required BPM value under this secondary classification combination, sequentially obtain the comprehensive feature vectors of the sub-music and calculate the similarity index with the reference comprehensive feature vector respectively to screen the candidate sub-music blocks, and generate AI music based on the GAN model and in combination with all the candidate sub-music blocks.

Citation Information

Patent Citations

  • Audio processing method and device, equipment and medium

    CN112382257A

  • Personalized music playing method and system and computer readable storage medium

    CN112464022A

  • Song creation method and related equipment

    CN113838445A

  • Specific emotion music generation method and system

    CN116030777A

  • Short video background music editing method based on deep learning

    CN116844562A