Method for generating electronic music album based on emotion matching

Through the emotional matching method, multi-dimensional feature analysis and emotion prediction model are used to generate electronic music albums, which solves the problem of relying on server-side templates and patterns in the existing technology, and achieves more accurate audio and video matching and richer user experience.

CN113656612BActive Publication Date: 2025-05-27COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110954182.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-19
Publication Date
2025-05-27
Estimated Expiration
2041-08-19

AI Technical Summary

Technical Problem

In the prior art, the generation of electronic music albums depends too much on the given templates and patterns provided by the server. The number of tracks selected is small and the type is single. Audio and video matching mainly depends on the user's manual filtering.

Method used

Using an emotion matching method, through multi-dimensional feature analysis of images and music, an emotion prediction model and an emotion matching model are used to generate an electronic music album. The specific steps include: multi-dimensional feature extraction based on images and music, building an emotion prediction model and an emotion matching model, obtaining emotion factors and emotion matching values ​​through these models, and finally generating an electronic music album.

Benefits of technology

It improves the accuracy of matching musical emotions and image emotions, reduces dependence on server-side templates and patterns, enhances user experience, and provides richer and more diverse music choices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113656612B_ABST
    Figure CN113656612B_ABST
Patent Text Reader

Abstract

An embodiment of the present disclosure provides a method for generating an electronic music album based on emotion matching. The method may include: obtaining a multi-dimensional image emotion prediction value by using an image emotion prediction model based on the multi-dimensional image features of an image; obtaining a multi-dimensional music emotion prediction value by using a music emotion prediction model based on the multi-dimensional music features of a music segment; obtaining a multi-dimensional image emotion factor prediction value by using a music-image emotion matching model based on the multi-dimensional image emotion prediction value; obtaining a multi-dimensional music emotion factor prediction value by using a music-image emotion matching model based on the multi-dimensional music emotion prediction value; and generating an electronic music album according to the multi-dimensional image emotion prediction value, the multi-dimensional music emotion prediction value, the multi-dimensional image emotion factor prediction value, and the multi-dimensional music emotion factor prediction value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method for intelligently generating an electronic photo album, and specifically, to a method for generating an electronic music photo album based on emotion matching. Background Art

[0002] With the rapid development of multimedia and Internet technologies, the form of manually making an electronic music photo album is gradually being replaced by an automated intelligent processing method. In related technologies, filters or music are selected according to the prompted scenarios, which simplifies the operation process of users.

[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in related technologies: The generation of electronic music photo albums relies too much on the given templates and patterns provided by the server side, the number of selectable tracks is small and the types are single, and the matching of audio and video mainly depends on the manual screening of images and music by users. Summary of the Invention

[0004] In view of this, the technical problem to be solved by the present disclosure is to provide a method for generating an electronic music photo album based on emotion matching, which solves the problem that the generation of electronic music photo albums in related technologies relies too much on the given templates and patterns provided by the server side.

[0005] To solve the above technical problem, a specific embodiment of the present disclosure provides a method for generating an electronic music photo album based on emotion matching, including: obtaining a multi-dimensional image emotion prediction value by using an image emotion prediction model based on the multi-dimensional image features of an image; obtaining a multi-dimensional music emotion prediction value by using a music emotion prediction model based on the multi-dimensional music features of a music segment; obtaining a multi-dimensional image emotion factor prediction value by using a music-image emotion matching model based on the multi-dimensional image emotion prediction value; obtaining a multi-dimensional music emotion factor prediction value by using a music-image emotion matching model based on the multi-dimensional music emotion prediction value; and generating an electronic music photo album according to the multi-dimensional image emotion prediction value, the multi-dimensional music emotion prediction value, the multi-dimensional image emotion factor prediction value, and the multi-dimensional music emotion factor prediction value.

[0006] Another aspect of the embodiments of the present disclosure provides an electronic device, including one or more processors and a storage device, wherein the storage device is used to store executable instructions, and when the executable instructions are executed by the processors, the method of the embodiments of the present disclosure is implemented.

[0007] Another aspect of the embodiments of the present disclosure provides a computer-readable storage medium, storing computer-executable instructions, and the instructions are used to implement the method of the embodiments of the present disclosure when executed by a processor.

[0008] Another aspect of the present disclosure provides a computer program, which includes computer-executable instructions that, when executed, are used to implement the method of the embodiments of the present disclosure.

[0009] According to the embodiments of the present disclosure, a mapping relationship between multiple sentiment factors and multi-dimensional sentiment adjectives is established, and then based on the features of images and music, a sentiment feature subset is constructed. Then, a machine learning algorithm is used to construct a sentiment prediction model based on the relationship between the sentiment feature subset and the sentiment factors. Then, the constructed sentiment prediction model and the above mapping relationship are used to complete the audio-visual matching. It can at least partially solve the problem that the generation of electronic music albums in the related art relies too much on the given templates and patterns provided by the server side, and thus can achieve the technical effect of improving the accuracy of matching music sentiment and image sentiment.

[0010] It should be understood that the above general description and the following specific embodiments are only exemplary and explanatory, and they do not limit the scope that the present disclosure intends to claim. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The following attached drawings are a part of the specification of the present disclosure, which illustrate the exemplary embodiments of the present disclosure. The attached drawings and the description of the specification are used together to explain the principles of the present disclosure.

[0012] Figure 1 It is a schematic overall flowchart of a method for generating an electronic music album based on sentiment matching provided by an embodiment of the present disclosure.

[0013] Figure 2 It is a schematic flowchart of training an image sentiment prediction model and a music sentiment prediction model provided by an embodiment of the present disclosure.

[0014] Figure 3 It is a schematic flowchart of obtaining multi-dimensional image sentiment prediction values by using an image sentiment prediction model based on multi-dimensional image features of an image provided by an embodiment of the present disclosure.

[0015] Figure 4 It is a schematic flowchart of obtaining multi-dimensional music sentiment prediction values by using a music sentiment prediction model based on multi-dimensional music features of a music segment provided by an embodiment of the present disclosure.

[0016] Figure 5 It is a schematic flowchart of training a music-image sentiment matching model provided by an embodiment of the present disclosure.

[0017] Figure 6 It is a schematic flowchart of generating an electronic music album provided by an embodiment of the present disclosure.

[0018] Figure 7 It is a schematic flowchart of dividing music into multiple music segments provided by an embodiment of the present disclosure. Detailed implementation manners

[0019] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the spirit of the content disclosed in the present disclosure will be clearly described below with reference to the accompanying drawings and detailed descriptions. After understanding the embodiments of the content of the present disclosure, any person skilled in the art can make changes and modifications to the techniques taught by the content of the present disclosure without departing from the spirit and scope of the content of the present disclosure.

[0020] The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure, but not to limit the present disclosure. In addition, the same or similar reference numerals of elements / components used in the accompanying drawings and embodiments are used to represent the same or similar parts.

[0021] Regarding the use of "first", "second",... etc. in this article, it does not particularly refer to the meaning of order or sequence, nor is it used to limit the present disclosure. It is only used to distinguish elements or operations described with the same technical terms.

[0022] Regarding the directional terms used in this article, such as: up, down, left, right, front or back, etc., they are only references to the directions in the accompanying drawings. Therefore, the directional terms used are for explanation and not for limiting the present creation.

[0023] Regarding the use of "comprising", "including", "having", "containing", etc. in this article, they are all open-ended terms, that is, they mean including but not limited to.

[0024] Regarding the use of "and / or" in this article, it includes any one or all combinations of the described things.

[0025] Regarding "a plurality" in this article, it includes "two" and "more than two"; regarding "multiple groups" in this article, it includes "two groups" and "more than two groups".

[0026] Regarding the terms "substantially", "about", etc. used in this article, they are used to modify any quantity or error that can vary slightly, but these slight variations or errors will not change their essence. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments or other values. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.

[0027] Figure 1 It is an overall schematic flowchart of a method for generating an electronic music album based on emotion matching provided for the embodiments of the present disclosure.

[0028] As Figure 1 shown, the method for generating an electronic music album based on emotion matching may include the following operations S101 to S105:

[0029] In operation S101, a multi-dimensional image emotion prediction value is obtained by using an image emotion prediction model based on multi-dimensional image features of an image.

[0030] In an embodiment of the present disclosure, the multi-dimensional image features include low-level features, high-level features, and key region features, etc. The low-level features mainly include dominant color, color moment, color contrast, gray-level co-occurrence matrix, Tamura features, etc.; the high-level features include aesthetic features and composition features, and the aesthetic features and composition features are introduced into the image feature extraction algorithm as high-level features, where the aesthetic features include texture complexity, color complexity, energy features, etc. of the image; the composition features include the rule of thirds, depth of field, dynamic features, etc. of the image; the key region feature algorithm is mainly extracted through three steps: image saliency calculation, salient region extraction, and salient region features. The image emotion prediction model is a pre-trained model. For example, the dimension of the image emotion value can be 18, and the 18-dimensional image emotion value can correspond to 18-dimensional image emotion adjectives, and the 18-dimensional image emotion adjectives can be as shown in Table 1 below.

[0031] Table 1

[0032] Serial number Emotional adjective Serial number Emotional adjective Serial number Emotional adjective 1 Comfortable 7 Anxious 13 Warm 2 Grand 8 Depressing 14 Romantic 3 Lonely 9 Dreamy 15 Hope 4 Depressed 10 Fear 16 Fresh 5 Happy 11 Sentimental 17 Warm and fragrant 6 Relaxed 12 Sunny 18 Lost

[0033] Then, in operation S102, a multi-dimensional music emotion prediction value is obtained by using a music emotion prediction model based on multi-dimensional music features of a music segment.

[0034] In an embodiment of the present disclosure, the multi-dimensional music features may include rhythm, spectral centroid, spectral contrast, MFCC (Mel Frequency Cepstrum Coefficient), chroma CQT (Constant Q Transform) spectrum, chroma spectral centroid, short-time average zero-crossing rate, spectral attenuation, etc. The music emotion prediction model is also a pre-trained model. For example, the dimension of the music emotion value can be 18, and the 18-dimensional music emotion value can correspond to 18-dimensional music emotion adjectives, and the 18-dimensional music emotion adjectives can also be the emotion adjectives shown in Table 1 above. The present disclosure can divide the music into multiple music segments by beat positions.

[0035] Next, in operation S103, a multi-dimensional image emotion factor prediction value is obtained by using a music-image emotion matching model based on the multi-dimensional image emotion prediction value.

[0036] In an embodiment of the present disclosure, the music-image emotion matching model is determined according to subjective experimental evaluation and factor analysis. The dimension of the multi-dimensional image emotion factor prediction value can be 4.

[0037] Second, in operation S104, based on the multi-dimensional music emotion prediction value, use the music image emotion matching model to obtain the multi-dimensional music emotion factor prediction value.

[0038] In an embodiment of the present disclosure, the dimension of the multi-dimensional music emotion factor prediction value is the same as the dimension of the multi-dimensional image emotion factor prediction value. If the dimension of the multi-dimensional image emotion factor prediction value is 4, then the dimension of the multi-dimensional music emotion factor prediction value is also 4.

[0039] Next, in operation S105, generate an electronic music album according to the multi-dimensional image emotion prediction value, the multi-dimensional music emotion prediction value, the multi-dimensional image emotion factor prediction value, and the multi-dimensional music emotion factor prediction value.

[0040] In an embodiment of the present disclosure, calculate the Euclidean distance (i.e., Euclidean distance) between the image and the music segment. Specifically, calculate the absolute value between the multi-dimensional image emotion prediction value corresponding to the image and the multi-dimensional music emotion prediction value corresponding to the music segment, calculate the absolute value between the multi-dimensional image emotion factor prediction value corresponding to the image and the multi-dimensional music emotion factor prediction value corresponding to the music segment, and then set weights for the two absolute values. If the sum of the products of each absolute value and the corresponding weight is smaller, it means that the emotional space between the image and the music segment is closer.

[0041] In an embodiment of the present disclosure, extract the underlying features, high-level features, and key region features of the image, etc., and extract features related to music attributes, such as energy, melody, timbre, rhythm, beat, harmony, etc. And match the image and the music segment based on the principle of the smallest Euclidean distance, that is, insert the image into the corresponding music segment, realize the automatic generation function of the electronic music album that more conforms to the human emotional cognitive system, improve the emotional matching accuracy in the production process of the music album, and enhance the user experience.

[0042] Figure 2 It is a schematic flowchart of a method for training an image emotion prediction model and a music emotion prediction model provided by an embodiment of the present disclosure.

[0043] As Figure 2 shown, before obtaining the multi-dimensional image emotion prediction value by using the image emotion prediction model based on the multi-dimensional image features of the image in operation S102, the method for generating an electronic music album based on emotion matching may further include the following operation S201:

[0044] In operation S201, train the image emotion prediction model based on the image feature training set and the image feature test set.

[0045] In an embodiment of the present disclosure, the image feature training set is used to debug the parameters of the image emotion prediction model, and the image feature test set is used to verify the image emotion prediction model.

[0046] Before obtaining the multi-dimensional music emotion prediction value using the music emotion prediction model based on the multi-dimensional music features in operation S102, the method for generating an electronic music album based on emotion matching may further include the following operation S202:

[0047] In operation S202, train a music emotion prediction model based on a music feature training set and a music feature test set.

[0048] In an embodiment of the present disclosure, the music feature training set is used to debug the parameters of the music emotion prediction model, and the music feature test set is used to test the music emotion prediction model.

[0049] In an embodiment of the present disclosure, pre-training an image emotion prediction model and a music emotion prediction model can improve the emotion matching accuracy in the process of making a music album.

[0050] In an embodiment of the present disclosure, operation S201 of training an image emotion prediction model based on an image feature training set and an image feature test set may include the following operations:

[0051] Randomly divide multiple standard image feature subsets into an image feature training set and an image feature test set according to a predetermined ratio; and

[0052] Train an image emotion prediction model corresponding to each image factor value according to the image feature training set, the image feature test set, and multiple image factor values respectively.

[0053] In the embodiments of the present disclosure, emotional adjectives that can reflect the emotions of an electronic music album can be widely collected. Then, 4 groups of music MV (Music Video) / MAD (Movie, Anime, Drama + Music video) videos can be designed as materials, and the content can include clips of movie and TV drama plots released after 2007 and original MV of popular music. The music languages can include Chinese, English, Japanese, Korean, etc. Multiple (not less than 15, such as 18, etc.) subjects can be selected to check the emotions stimulated by the audio-visual content. After high-frequency word screening and principal base analysis, multiple (such as 18) emotional adjectives can be finally retained. The finally retained emotional adjectives can be used for multi-dimensional emotion recognition and matching, and Table 1 above can be referred to. A subjective evaluation experiment can be carried out on 18 image emotional adjectives as shown in Table 1 above to obtain 18-dimensional image emotion values. For example, each image emotion value is represented by a scale of 1-5 in a psychological scale, where 1 represents no such emotion and 5 represents very strong. After normalization, this value will be standardized to between 0 and 1, and the corresponding image emotion value can be obtained. Factor analysis method is used to establish the mapping relationship between the image factor value and the multi-dimensional image emotion value from the perspective of emotion description. The multi-dimensional image factor value can be obtained from the multi-dimensional image emotion value. For example, the 5-dimensional image factor value can be obtained from the 18-dimensional image emotion value; the multi-dimensional image emotion value can be obtained from the multi-dimensional image factor value. For example, the 18-dimensional image emotion value can be obtained from the 5-dimensional image factor value.

[0054] The present disclosure can refer to the establishment standards and principles of IAPS (International Affective Picture System) and CAPS (Chinese Affective Picture System), and balance from multiple aspects such as color, people, animals, buildings, natural scenery, etc., to establish a number of (such as 700 to 1000) image databases covering all original emotions. Each standard image in the image database reflects the image emotion value corresponding to each image emotional adjective. The features of all standard images in the image database are divided into multiple standard image feature subsets, and the standard image feature subsets correspond to the image factor values. Among them, the image features in the standard image feature subsets are the image features after dimensionality reduction. Since there is a mapping relationship between the image factor value and the multi-dimensional image emotion value, the mapping relationship between the standard image feature subset and the image factor value can be obtained.

[0055] In the embodiments of the present disclosure, multiple standard image feature subsets can be randomly divided into an image feature training set and an image feature test set according to a ratio of 8:2.

[0056] In the embodiments of the present disclosure, operation S202 trains a music emotion prediction model based on the music feature training set and the music feature test set, and may include the following operations:

[0057] Randomly divide multiple subsets of standard music features into a music feature training set and a music feature test set according to a predetermined ratio; and

[0058] Train a music emotion prediction model corresponding to each music factor value respectively according to the music feature training set, the music feature test set, and multiple music factor values.

[0059] In the embodiments of the present disclosure, emotional adjectives that can reflect the emotions of an electronic music album can be widely collected, and then 4 groups of music MV (Music Video) / MAD (Movie, Anime, Drama + Music video) videos can be designed as materials. The content can include clips of movie and TV drama plots produced after 2007 and original MV of pop music. The music languages can include Chinese, English, Japanese, Korean, etc. Multiple (not less than 15, such as 18) subjects can be selected to check the emotions inspired by the audio-visual content. After high-frequency word screening and principal component analysis, multiple (such as 18) emotional adjectives can be finally retained. The finally retained emotional adjectives can be used for multi-dimensional emotion recognition and matching, and Table 1 above can be referred to. A subjective evaluation experiment can be carried out on the 18 music emotion adjectives as shown in Table 1 above to obtain 18-dimensional music emotion values. For example, each music emotion value is represented by a scale of 1-5 in a psychological scale, where 1 represents no such emotion and 5 represents very strong. After normalization, this value will be normalized to between 0 and 1, and the corresponding music emotion value can be obtained. The factor analysis method is used to establish the mapping relationship between the music factor value and the multi-dimensional music emotion value from the perspective of emotion description. The multi-dimensional music factor value can be obtained from the multi-dimensional music emotion value. For example, the 4-dimensional music factor value can be obtained from the 18-dimensional music emotion value; the multi-dimensional music emotion value can be obtained from the multi-dimensional music factor value. For example, the 18-dimensional music emotion value can be obtained from the 4-dimensional music factor value.

[0060] The present disclosure can widely collect diverse types of music materials covering all original emotions and establish a music database containing at least multiple (such as 700 to 1000) pieces of music. Each standard music segment in the music database reflects the music emotion value corresponding to each music emotion adjective. The features of all standard music segments in the music database are divided into multiple subsets of standard music features, and each subset of standard music features corresponds to a music factor value. Since there is a mapping relationship between the music factor value and the multi-dimensional music emotion value, the mapping relationship between the subset of standard music features and the music factor value can be obtained.

[0061] In the embodiments of the present disclosure, multiple subsets of standard music features can be randomly divided into a music feature training set and a music feature test set according to a ratio of 8:2.

[0062] In an embodiment of the present disclosure, by separately training an image emotion prediction model corresponding to each image factor value and a music emotion prediction model corresponding to each music factor value, the emotion matching accuracy in the process of making a music album can be improved, enhancing the user experience.

[0063] Figure 3 FIG. is a schematic flowchart of obtaining a multi-dimensional image emotion prediction value by using an image emotion prediction model based on multi-dimensional image features of an image provided by an embodiment of the present disclosure.

[0064] As Figure 3 shown, operation S101 of obtaining a multi-dimensional image emotion prediction value by using an image emotion prediction model based on multi-dimensional image features of an image may include the following operations S1011 to S1014:

[0065] In operation S1011, multi-dimensional image features related to emotion description are extracted from different levels of the image.

[0066] In an embodiment of the present disclosure, multi-dimensional image features related to emotion description are extracted from the bottom layer, high layer, key regions, etc. of the image. The multi-dimensional image features may include, for example, bottom layer features, high layer features, and key region features. Among them, the bottom layer features mainly include the main color, color moment, color contrast, gray level co-occurrence matrix, Tamura features, etc.; aesthetic features and composition features can be introduced into the image feature extraction algorithm as high layer features. Among them, the aesthetic features may include the texture complexity, color complexity, energy features, etc. of the image, and the composition features may include the rule of thirds, depth of field, dynamic features, etc. of the image; the key region features are mainly extracted through three steps: image saliency calculation, salient region extraction, and salient region features.

[0067] Secondly, in operation S1012, the multi-dimensional image features are dimensionally reduced for each image factor value to obtain an image emotion feature subset corresponding to the image factor value.

[0068] In an embodiment of the present disclosure, the multi-dimensional image features are screened and dimensionally reduced based on Lasso (Least absolute shrinkage and selection operator) regression and RF (random forest) feature importance to obtain an image emotion feature subset corresponding to the image factor value. Dimensionally reducing the multi-dimensional image features can improve the generalization ability of the model and reduce the difficulty of learning.

[0069] Then, in operation S1013, an image factor prediction value is obtained by using an image emotion prediction model corresponding to each image factor value based on the image emotion feature subset corresponding to each image factor value.

[0070] In an embodiment of the present disclosure, each image factor value corresponds to an image emotion feature subset, and each image factor value corresponds to an image emotion prediction model. The image emotion feature subset corresponding to each image factor value is input into the image emotion prediction model corresponding to the image factor value to obtain the image factor prediction value. The MSE (mean squared error) and MAE (mean absolute error) can be used as the loss function of the image emotion prediction model.

[0071] In an embodiment of the present disclosure, the image factor prediction value can be represented by . For example, if there are 5 image factor prediction values, then m is 5. The image factor prediction value can explain more than 90% of the information of multiple (e.g., 18) image emotion adjectives.

[0072] Next, in operation S1014, based on all the image factor prediction values, the multi-dimensional image emotion prediction value is obtained by using the inverse transformation formula.

[0073] In an embodiment of the present disclosure, after obtaining the image factor prediction value, the image emotion prediction value can be obtained by using the following inverse transformation formula.

[0074]

[0075] In the above formula, is the image emotion prediction value. The image emotion prediction value represents the excitation degree of the image predicted by the image emotion prediction model on 18 emotion adjectives. The larger the value, the stronger the emotion corresponding to the emotion adjective.

[0076] In an embodiment of the present disclosure, the generalization ability of the model can be improved and the learning difficulty can be reduced. It can also improve the emotion matching accuracy in the process of making music albums and enhance the user experience.

[0077] Figure 4 FIG. is a schematic flowchart of obtaining a multi-dimensional music emotion prediction value by using a music emotion prediction model based on multi-dimensional music features of a music segment provided by an embodiment of the present disclosure.

[0078] As Figure 4 shown, the operation S102 of obtaining the multi-dimensional music emotion prediction value by using the music emotion prediction model based on the multi-dimensional music features of the music segment may include the following operations S1021 to S1024:

[0079] In operation S1021, multi-dimensional music features of the music segment related to emotion expression are extracted from multiple perspectives.

[0080] In the embodiments of the present disclosure, the multi-dimensional music features may include rhythm, spectral centroid, spectral contrast, MFCC (Mel Frequency Cepstrum Coefficient), chromatic CQT (Constant Q Transform) spectrum, chromatic spectral centroid, short-time average zero-crossing rate, spectral attenuation, and the like.

[0081] Secondly, in operation S1022, the multi-dimensional music features are dimensionally reduced for each music factor value to obtain a music emotion feature subset corresponding to the music factor value.

[0082] In the embodiments of the present disclosure, based on Lasso (Least absolute shrinkage and selection operator) regression and RF (random forest) feature importance, the multi-dimensional music features are screened and dimensionally reduced to obtain a music emotion feature subset corresponding to the music factor value. Dimensionally reducing the multi-dimensional music features can improve the generalization ability of the model and reduce the difficulty of learning.

[0083] Next, in operation S1023, based on the music emotion feature subset corresponding to each music factor value, the music factor prediction value is obtained by using the music emotion prediction model corresponding to the music factor value.

[0084] In the embodiments of the present disclosure, each music factor value corresponds to a music emotion feature subset, and each music factor value corresponds to a music emotion prediction model. The music emotion feature subset corresponding to each music factor value is input into the music emotion prediction model corresponding to the music factor value to obtain the music factor prediction value. MSE (Mean Squared Error) and MAE (Mean Absolute Error) can be used as the loss functions of the music emotion prediction model.

[0085] In the embodiments of the present disclosure, the music factor prediction value can be represented by For example, if there are 4 music factor prediction values, then n is 4. The music factor prediction value can explain more than 90% of the information of multiple (e.g., 18) music emotion adjectives.

[0086] Then, in operation S1024, based on all the music factor prediction values, the multi-dimensional music emotion prediction value is obtained by using the inverse transformation formula.

[0087] In the embodiments of the present disclosure, after obtaining the music factor prediction value, the music emotion prediction value can be obtained by using the following inverse transformation formula.

[0088]

[0089] Among them, in the above formula, represents the music emotion prediction value. The music emotion prediction value obtained from the music emotion prediction model represents the excitation degree of the music on 18 emotion adjectives. The larger the value, the stronger the emotion corresponding to the emotion adjective.

[0090] In the embodiments of the present disclosure, the generalization ability of the model can be improved, and the learning difficulty can be reduced. Moreover, the emotion matching accuracy in the process of making a music photo album can be enhanced, and the user experience can be improved.

[0091] Figure 5 FIG. is a schematic flowchart of a method for training a music image emotion matching model provided by an embodiment of the present disclosure.

[0092] As Figure 5 shown, before obtaining the multi-dimensional image emotion factor prediction value by using the music image emotion matching model based on the multi-dimensional image emotion prediction value in operation S103, the method for generating an electronic music photo album based on emotion matching may further include the following operation S501:

[0093] In operation S501, a factor analysis method is used to construct a music image emotion matching model, where the music image emotion matching model is used to describe the mapping relationship between the emotion prediction value and the emotion factor prediction value.

[0094] In the embodiments of the present disclosure, using the factor analysis method to construct a music image emotion matching model can improve the emotion matching accuracy in the process of making a music photo album.

[0095] In the embodiments of the present disclosure, the factor analysis method can evaluate the activation degree of multi-dimensional (e.g., 18-dimensional) emotion adjectives and the emotion matching degree of music images. The activation degree can be represented by 1-5 points based on a psychological scale, where 1 point means not feeling the emotion, and 5 points means the emotion is very strong. The matching degree is also represented by 1-5 points based on a psychological scale, where 1 point means not matching, and 5 points means very matching. Retain the data with a matching degree of 3 or more, and take the average value of the evaluation results of all subjects in these data. After normalization, the emotion evaluation result of the user for the "image + music" data set is obtained, that is, the music image emotion factor prediction value corresponding to the multi-dimensional emotion adjective If there are 18 multi-dimensional emotion adjectives, then the value of j is 18. At this time, the music image emotion matching model f l MTCH , l = 1, 2, 3,..., s can be expressed as follows:

[0096]

[0097] Among them, in the above formula, there are 4 music image factor prediction values, that is, l is 4. When For the image emotion prediction value of the image (i.e., ), the multi-dimensional image emotion factor prediction value f l MTCH_im can be calculated, where l = 1, 2, 3, 4. At this time, f l MTCH = f l MTCH_im , l = 1, 2, 3, 4. Similarly, when is the music emotion prediction value of the music (i.e., ), the multi-dimensional music emotion factor prediction value f l MTCH_m can be calculated, where l = 1, 2, 3, 4. At this time, f l MTCH = f l MTCH_m , l = 1, 2, 3, 4.

[0098] In the embodiments of the present disclosure, compared with separately constructing models for images and music respectively, the multi-dimensional image emotion factor prediction value and the multi-dimensional music emotion factor prediction value with the same dimension can be obtained, and the emotion matching accuracy in the process of making a music album can be improved, enhancing the user experience.

[0099] Figure 6 FIG. is a schematic flowchart of an electronic music album generation provided by an embodiment of the present disclosure.

[0100] As Figure 6 shown, the operation S105 of generating an electronic music album according to the multi-dimensional image emotion prediction value, the multi-dimensional music emotion prediction value, the multi-dimensional image emotion factor prediction value, and the multi-dimensional music emotion factor prediction value may include the following operations S1051 to S1053:

[0101] In operation S1051, a first Euclidean distance is obtained according to the multi-dimensional image emotion prediction value and the multi-dimensional music emotion prediction value using the Euclidean distance calculation formula.

[0102] In the embodiments of the present disclosure, assuming there are two points x and y in a k-dimensional space, and their coordinates in any dimension are represented by x i and y i , where i = 1, 2,... k, then the calculation method of the Euclidean distance dist(X, Y) between these two points x and y is:

[0103]

[0104] In the embodiments of the present disclosure, if there are 18-dimensional image emotion prediction values and 18-dimensional music emotion prediction values, the calculation formula of the first Euclidean distance dist1 can be as follows:

[0105]

[0106] Among them, in the above formula, represents the music emotion prediction value; represents the image emotion prediction value; j represents the dimension serial number of the image emotion prediction value and the music emotion prediction value.

[0107] Secondly, in operation S1052, according to the multi-dimensional image emotion factor prediction value and the multi-dimensional music emotion factor prediction value, the second Euclidean distance is obtained by using the Euclidean distance calculation formula.

[0108] In the embodiments of the present disclosure, assuming there are 4-dimensional image emotion factor prediction values and 4-dimensional music emotion factor prediction values, the calculation formula of the second Euclidean distance dist2 can be as follows:

[0109]

[0110] Among them, in the above formula, f l MTCH_m represents the music emotion factor prediction value; f l MTCH_im represents the image emotion factor prediction value; l represents the dimension serial number of the image emotion factor prediction value and the music emotion factor prediction value.

[0111] Next, in operation S1053, according to the first Euclidean distance, the second Euclidean distance and the predetermined weight, the music-image emotion matching value between each image and each music segment is obtained.

[0112] In the embodiments of the present disclosure, the calculation formula of the music-image emotion matching value dist can be represented by the following formula:

[0113] dist = αdist1 + (1 - α)dist2

[0114] Among them, in the above formula, α represents the weight, and 0 ≤ α ≤ 1.

[0115] Then, in operation S1054, according to the music-image emotion matching value, the image most matching each music segment is determined.

[0116] In the embodiments of the present disclosure, the smaller the music-image emotion matching value, the more matching the image and the music segment are, so that the image most matching each music segment can be found.

[0117] Next, in operation S1055, each image is inserted at the position corresponding to the most matching music segment.

[0118] In an embodiment of the present disclosure, the most matching image is inserted at the corresponding music segment. For example, the most matching image is inserted at the beat position of each music segment. After inserting each image at the corresponding most matching music segment in operation S1055, operation S105 may further include the following operation: setting a transition effect for the inserted image. The transition effects include wipe, dissolve, fade in and fade out, fly in and fly out, page curl, etc.

[0119] In an embodiment of the present disclosure, adjusting the weights of the first Euclidean distance and the second Euclidean distance can take into account the advantages of matching based on emotional adjectives and factor values, so that the matched image and music segment are not only the closest in terms of the distance of a single emotional adjective, but also the most similar in terms of the type of emotional expression, thereby improving the emotional matching accuracy and enhancing the user experience.

[0120] Figure 7 It is a schematic flowchart of dividing music into multiple music segments provided by an embodiment of the present disclosure.

[0121] As Figure 7 shown, before obtaining the multi-dimensional music emotion prediction value by using the music emotion prediction model based on the multi-dimensional music features of the music segment in operation S102, the method for generating an electronic music album based on emotion matching may further include the following operation S701:

[0122] In operation S701, the music is divided into multiple music segments by using a beat detection algorithm.

[0123] In an embodiment of the present disclosure, the beat detection algorithm includes a deterministic algorithm and a learning algorithm. The music is divided into multiple music segments so that the image can be accurately matched with the music segment, and the image is inserted into the most matching music segment.

[0124] In an embodiment of the present disclosure, operation S701 using the beat detection algorithm to divide the music into multiple music segments may include the following operations:

[0125] The music is framed at a unified sampling rate to obtain multiple audio frames. Among them, according to the music playing order, the audio frames have fixed serial numbers.

[0126] The constant Q transform (CQT) is used to obtain the constant Q transform value (CQT value) corresponding to each audio frame. Among them, the CQT value is similar to the spectrum value.

[0127] The audio frames are screened according to the constant Q transform values corresponding to all audio frames. Among them, only the audio frames with CQT values greater than a predetermined value are retained.

[0128] Multiple note start sequences of the music are obtained according to the screened audio frames. Among them, the screened audio frames correspond to the note start sequences of the music.

[0129] After performing Gaussian filtering and smoothing on the obtained multiple note start sequences, extract the peaks of the multiple note start sequences.

[0130] Obtain the tempo of the music based on the peaks. Specifically, the tempo of the music can be obtained by dividing the time from the first peak to the last peak by the number of peaks.

[0131] Based on the tempo, re-divide the time interval of each beat from the last peak of the music backwards. Here, the tempo refers to the time interval of each beat.

[0132] Divide the music into multiple music segments according to the time interval of each beat and the rhythm factor. Among them, the product of the time interval and the rhythm factor is the duration of the music segment. For example, the rhythm factor of a three-beat music is 3, and the rhythm factor of a four-beat music is 4.

[0133] In the embodiments of the present disclosure, obtaining the tempo of the music based on the peaks may include the following operations:

[0134] Perform amplitude sampling on the note start sequence to determine the amplitude threshold of the beat points;

[0135] Filter out the peaks smaller than the amplitude threshold according to the determined amplitude threshold to obtain valid peaks;

[0136] Place the first time interval between adjacent valid peaks in an array;

[0137] Remove the second time intervals that do not conform to the song speed from the array. At this time, the remaining third time intervals are in the array; among them, the speed of the song is generally between 60 beats per minute and 220 beats per minute, that is, the tempo is approximately between 0.27 and 1 second. If the first time interval is not between 0.27 seconds and 1 second, it is regarded as a second time interval to be removed;

[0138] Retain one decimal place of the third time interval to obtain the fourth time interval;

[0139] Obtain the mode of the fourth time interval;

[0140] Retain the third time interval according to the mode; among them, the third time intervals located in the interval [mode - 0.05, mode + 0.05] can be retained;

[0141] Redetermine the peaks according to the third time interval so as to obtain the tempo of the music based on the peaks.

[0142] In the embodiments of the present disclosure, the music segments obtained through the above algorithm are more in line with the distribution law of the music signal and closer to the human auditory system, laying a good foundation for the accurate matching of subsequent images and music segments, thereby improving the emotional matching accuracy in the production process of music albums and enhancing the user experience.

[0143] In the embodiments of the present disclosure, the method for generating an electronic music album based on emotional matching may further include the following operations:

[0144] Play the generated electronic music album according to the preset settings.

[0145] In the embodiments of the present disclosure, the preset settings may include image display settings, playback interface settings, etc. Specifically, the MATLAB GUI tool can be used to complete the construction and setting of the electronic music album interface.

[0146] The above are only illustrative specific embodiments of the present disclosure. Without departing from the concept and principles of the present disclosure, any equivalent changes and modifications made by any person skilled in the art shall fall within the scope of protection of the present disclosure.

Claims

1. A method for generating an electronic music album based on emotion matching, including: obtaining a multi-dimensional image emotion prediction value by using an image emotion prediction model based on the multi-dimensional image features of an image; obtaining a multi-dimensional music emotion prediction value by using a music emotion prediction model based on the multi-dimensional music features of a music segment; constructing a music-image emotion matching model by using factor analysis, wherein the music-image emotion matching model is used to describe the mapping relationship between the emotion prediction value and the emotion factor prediction value; obtaining a multi-dimensional image emotion factor prediction value by using the music-image emotion matching model based on the multi-dimensional image emotion prediction value, wherein the music-image emotion matching model is composed of multi-dimensional image emotion factor prediction values, and each dimension of image emotion factor prediction value is composed of multi-dimensional image emotion prediction values; obtaining a multi-dimensional music emotion factor prediction value by using the music-image emotion matching model based on the multi-dimensional music emotion prediction value, wherein the music-image emotion matching model is composed of multi-dimensional music emotion factor prediction values, and each dimension of music emotion factor prediction value is composed of multi-dimensional music emotion prediction values; and generating an electronic music album according to the multi-dimensional image emotion prediction value, the multi-dimensional music emotion prediction value, the multi-dimensional image emotion factor prediction value, and the multi-dimensional music emotion factor prediction value.

2. The method according to claim 1, further including: before obtaining a multi-dimensional image emotion prediction value by using an image emotion prediction model based on the multi-dimensional image features of an image, training the image emotion prediction model based on an image feature training set and an image feature test set, before obtaining a multi-dimensional music emotion prediction value by using a music emotion prediction model based on the multi-dimensional music features of a music segment, training the music emotion prediction model based on a music feature training set and a music feature test set.

3. The method according to claim 2, wherein, training the image emotion prediction model based on an image feature training set and an image feature test set includes: randomly dividing a plurality of standard image feature subsets into the image feature training set and the image feature test set according to a predetermined ratio; and training the image emotion prediction model corresponding to each image factor value respectively according to the image feature training set, the image feature test set, and a plurality of image factor values, training the music emotion prediction model based on a music feature training set and a music feature test set includes: randomly dividing a plurality of standard music feature subsets into the music feature training set and the music feature test set according to a predetermined ratio; and training the music emotion prediction model corresponding to each music factor value respectively according to the music feature training set, the music feature test set, and a plurality of music factor values.

4. The method according to claim 1, wherein, obtaining a multi-dimensional image emotion prediction value by using an image emotion prediction model based on the multi-dimensional image features of an image includes: extracting the multi-dimensional image features related to emotion description from different levels of the image; dimensionality reduction is performed on the multi-dimensional image features for each image factor value to obtain an image emotion feature subset corresponding to the image factor value; Obtain an image factor prediction value by using the image emotion prediction model corresponding to the image factor value based on the subset of image emotion features corresponding to each image factor value; and Obtain a multi-dimensional image emotion prediction value by using an inverse transformation formula based on all the image factor prediction values.

5. The method according to claim 1, wherein, Obtaining a multi-dimensional music emotion prediction value by using a music emotion prediction model based on the multi-dimensional music features of a music segment includes: Extracting the multi-dimensional music features of the music segment related to emotion expression from multiple perspectives; Reducing the dimension of the multi-dimensional music features for each music factor value to obtain a subset of music emotion features corresponding to the music factor value; Obtain a music factor prediction value by using the music emotion prediction model corresponding to the music factor value based on the subset of music emotion features corresponding to each music factor value; and Obtain a multi-dimensional music emotion prediction value by using an inverse transformation formula based on all the music factor prediction values.

6. The method according to claim 1, wherein, Generating an electronic music album according to the multi-dimensional image emotion prediction value, the multi-dimensional music emotion prediction value, the multi-dimensional image emotion factor prediction value and the multi-dimensional music emotion factor prediction value includes: Obtaining a first Euclidean distance by using a Euclidean distance calculation formula according to the multi-dimensional image emotion prediction value and the multi-dimensional music emotion prediction value; Obtaining a second Euclidean distance by using a Euclidean distance calculation formula according to the multi-dimensional image emotion factor prediction value and the multi-dimensional music emotion factor prediction value; Obtaining a music-image emotion matching value between each image and each music segment according to the first Euclidean distance, the second Euclidean distance and a predetermined weight; Determining the image most matching each music segment according to the music-image emotion matching value; and Inserting each image at the position of the most matching music segment.

7. The method according to claim 6, wherein, Generating an electronic music album according to the multi-dimensional image emotion prediction value, the multi-dimensional music emotion prediction value, the multi-dimensional image emotion factor prediction value and the multi-dimensional music emotion factor prediction value further includes: After inserting each image at the position of the most matching music segment, Setting a transition effect for the inserted image.

8. The method according to claim 1, further includes: Before obtaining a multi-dimensional music emotion prediction value by using a music emotion prediction model based on the multi-dimensional music features of a music segment, Dividing the music into multiple music segments by using a beat detection algorithm.

9. The method according to claim 8, wherein, Dividing the music into multiple music segments by using a beat detection algorithm includes: Framing the music at a unified sampling rate to obtain multiple audio frames; Obtaining a constant Q transform value corresponding to each audio frame by using a constant Q transform; Screening audio frames according to the constant Q transform values corresponding to all the audio frames; Obtaining multiple note start sequences of the music according to the screened audio frames; Performing Gaussian filtering and smoothing processing on the obtained multiple note start sequences, and then extracting the peaks of the multiple note start sequences; Obtaining the tempo of the music according to the peaks; According to the tempo, starting from the last peak of the music and moving backward, re-dividing the time interval of each beat; and Divide the music into multiple music segments according to the time intervals and rhythm factors of each re-divided beat.

Citation Information

Patent Citations

  • Music adding method and device

    CN108153831A

  • Vehicle following state driver emotion dynamic feature extraction and identification method

    CN109858738A