Chord prediction method and device, computer equipment and storage medium
By obtaining the emotional parameters of the lyrics and using the harmony and pitch prediction model to generate matching emotional harmony progress and pitch probability sequences, the problem of chord design mismatch with user emotions is solved, and the quality of music generation and user experience is improved.
Patent Information
- Application Number
- CN202510243512.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-27
AI Technical Summary
The chord designs that generate music in the prior art are likely to mismatch the emotions that users need to express, resulting in poor user experience.
By obtaining the emotional parameters of the target lyrics in at least one emotional dimension, input the harmony to perform a prediction model, and generate the target harmony; then, input the target harmony to perform a pitch prediction model, obtain the pitch probability sequence of each chord, and determine the recommended tones and avoidance of each chord based on this.
It effectively solves the problem that the generation of music harmony is mismatched with the emotions that users need to express, improves the quality of music generation and user experience, and reduces out-of-tune and discord.
Smart Images

Figure CN120048233A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of music generation, and in particular, to a chord prediction method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of artificial intelligence technology, a variety of applications have been derived, including automatically arranging music according to the lyrics input by users. A relatively key part in music arrangement is to determine the harmonic progression and the pitch of each note included in each chord in the harmonic progression, which will affect the overall expression of the music. However, due to the randomness of the artificial intelligence-generated content technology, the chord design in the generated music may not conform to the music theory rules. As a result, users may think that the generated music is out of tune, disharmonious, and does not match the emotions that the users need to express, bringing a bad experience to the users. Summary of the Invention
[0003] The purpose of this application aims to at least solve one of the above technical defects, especially the defect that the chord design of the generated music in the prior art is prone to problems such as not matching the emotions that the users need to express.
[0004] In a first aspect, an embodiment of this application provides a chord prediction method, including:
[0005] Obtain the emotional parameters of the target lyrics in at least one emotional dimension;
[0006] Input the emotional parameters into the harmonic progression prediction model to obtain the target harmonic progression;
[0007] Input the target harmonic progression into the pitch prediction model to obtain the pitch probability sequence of each chord in the target harmonic progression;
[0008] Determine the recommended notes and avoidance notes corresponding to each chord in the target harmonic progression according to the pitch probability sequence.
[0009] In one embodiment, the training process of the harmonic progression prediction model includes:
[0010] Obtain a first training set with first annotation information; the first training set includes multiple first training harmonic progressions, and the first annotation information includes the emotional parameters of the first training harmonic progressions in at least one emotional dimension;
[0011] Unify the key of each first training harmonic progression in the first training set to the target key;
[0012] Use the first training set with unified key to train the initial harmonic progression prediction model to obtain the harmonic progression prediction model.
[0013] In one embodiment, obtaining the first training set with first annotation information includes:
[0014] Obtain the first training harmonic progression and its corresponding initial set of emotional parameters; the initial set of emotional parameters includes the initial emotional parameters annotated by multiple annotators for the first training harmonic progression in at least one emotional dimension;
[0015] For any emotional dimension, filter out invalid annotations in the initial emotional parameters, and obtain the emotional parameter corresponding to this emotional dimension according to the first average value of the initial emotional parameters after filtering.
[0016] In one embodiment, for any emotional dimension, filtering out invalid annotations in the initial emotional parameters includes:
[0017] For any emotional dimension, determine the second average value and standard deviation of all initial emotional parameters under this emotional dimension;
[0018] If the absolute value of the difference between the initial emotional parameter and the second average value is greater than the standard deviation, then determine this initial emotional parameter as an invalid annotation.
[0019] In one embodiment, the number of the first training harmonic progressions of the first training set in each emotional classification quadrant is greater than a preset lower limit, and the difference in the number of the first training harmonic progressions in different emotional classification quadrants is less than a deviation threshold.
[0020] In one embodiment, the training process of the pitch prediction model includes:
[0021] Obtain a second training set with second annotation information; the second training set includes multiple second training harmonic progressions with the tonality unified to the target tonality, and the second annotation information includes the notes contained in each chord of the second training harmonic progression;
[0022] Use the second training set to train the initial pitch prediction model to obtain the pitch prediction model.
[0023] In one embodiment, according to the pitch probability sequence, determining the recommended notes corresponding to each chord in the target harmonic progression includes:
[0024] Determine the in-chord notes and out-of-chord notes according to the corresponding chord type;
[0025] Arrange them in descending order of probability, and select the in-chord notes among the top three notes with the highest probability in the pitch probability sequence and the out-of-chord notes among the top five notes with the highest probability as the recommended notes.
[0026] In one embodiment, according to the pitch probability sequence, determining the avoidance notes corresponding to each chord in the target harmonic progression includes:
[0027] Determine the out-of-key notes according to the target tonality;
[0028] Determine non-chord tones according to the corresponding chord types;
[0029] Take the non-chord tones among the two notes with the lowest probabilities in the non-chord tones and the note sequence of pitch probabilities except for the non-chord tones as the avoidance tones.
[0030] In a second aspect, an embodiment of the present application provides a chord prediction device, including:
[0031] A data acquisition module, configured to acquire emotional parameters of the target lyrics in at least one emotional dimension;
[0032] A harmonic progression prediction module, configured to input the emotional parameters into a harmonic progression prediction model to obtain a target harmonic progression;
[0033] A pitch probability prediction module, configured to input the target harmonic progression into a pitch prediction model to obtain a pitch probability sequence of each chord in the target harmonic progression;
[0034] A pitch range determination module, configured to determine recommended tones and avoidance tones corresponding to each chord in the target harmonic progression according to the pitch probability sequence.
[0035] In a third aspect, an embodiment of the present application provides a computer device, including one or more processors and a memory. When computer-readable instructions stored in the memory are executed by one or more processors, the steps of the chord prediction method in any of the above embodiments are executed.
[0036] In a fourth aspect, an embodiment of the present application provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to execute the steps of the chord prediction method in any of the above embodiments.
[0037] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0038] This solution first obtains the emotional parameters of the target lyrics input by the user in at least one emotional dimension. Then, the obtained emotional parameters are input into a specially trained harmony progression prediction model to generate a target harmony progression that matches the emotion of the target lyrics. The predicted target harmony progression is input into another specially trained pitch prediction model to obtain the pitch probability sequence corresponding to each chord in the target harmony progression. Finally, based on the pitch probability sequence, recommended pitches and avoided pitches are determined for each chord. This solution learns the music theory rules contained in existing music works through machine learning, effectively solving the problem that the generated music harmony progression does not match the emotion that the user needs to express. At the same time, it provides a reference basis for solving the problems of out-of-tune and disharmony in artificial intelligence-generated music. Based on this solution, the works generated by artificial intelligence can meet higher quality requirements in terms of emotion expression, music rationality, etc., improving the feasibility and practicality of artificial intelligence music creation. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 It is a schematic flowchart of a chord prediction method provided by an embodiment of the present application;
[0041] Figure 2 It is a schematic block diagram of a chord prediction device provided by an embodiment of the present application;
[0042] Figure 3 It is an internal structure diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0044] Please refer to Figure 1 , an embodiment of the present application provides a chord prediction method, including steps S102 to S108.
[0045] S102, obtain the emotional parameters of the target lyrics in at least one emotional dimension.
[0046] It can be understood that the target lyrics are the lyrics given by the user during automatic music arrangement. The emotional dimension is a dimension used to quantify the emotional characteristics of a given input, which provides a way to map subjective emotional experiences to objective numerical values, and the emotional parameter is the quantization result evaluated using the quantization method corresponding to the emotional dimension. To expand the richness of emotional expression, two or more emotional dimensions can be used. Commonly used emotional dimensions can be two emotional dimensions related to the VA (Valence-Arousal) emotion model, namely the Valence dimension and the Arousal dimension. The emotional parameter corresponding to the Valence dimension indicates the degree of positiveness of the emotion. The higher the Valence, the more positive the corresponding emotion; the lower the Valence, the more negative the corresponding emotion. The emotional parameter corresponding to the Arousal dimension indicates the intensity or excitement level of the emotion. The higher the Arousal, the more excited and agitated the emotion; the lower the Arousal, the calmer and more peaceful the emotion. In the following, the emotional dimensions of Valence and Arousal will be used as examples for illustration.
[0047] Obtaining the emotional parameters of the target lyrics in at least one emotional dimension can be done manually, but it is time-consuming and laborious. Therefore, generally, rule dictionaries, deep learning models, etc. are used for automatic evaluation. Taking the rule dictionary solution as an example, specifically, first, a certain number of emotional seed words are obtained, and the emotional parameters of the emotional seed words in the selected emotional dimension are evaluated. Then, Word2Vec is used to calculate the similarity between each non-seed word and the emotional seed words. Based on the highest similarity as the judgment basis, words with a similarity higher than the similarity threshold are also added to the emotional dictionary, and the emotional parameters of the most similar emotional seed words are assigned to the non-emotional seed words newly added to the emotional dictionary. The emotional parameters of non-emotional seed words can also be fine-tuned manually. After constructing the emotional dictionary, the target lyrics are segmented. Word2Vec is used to calculate the similarity between each segmented word and the emotional words in the emotional dictionary. If there is an emotional word with a similarity greater than the matching threshold, the emotional parameter of the corresponding emotional word is assigned to the corresponding segmented word. All the segmented words in the target lyrics that are assigned emotional parameters are regarded as emotional segmented words. According to the preset correspondence relationship between the part of speech of each emotional segmented word and the role of emotional expression, a weight is assigned to each emotional segmented word (the sum of all weights is 1). For example, nouns are assigned a higher weight. Based on the weights of each emotional segmented word and the assigned emotional parameters, a weighted sum is calculated to obtain the emotional parameters of the target lyrics in at least one emotional dimension.
[0048] S104, Input the emotional parameters into the harmony progression prediction model to obtain the target harmony progression.
[0049] It can be understood that the harmonic progression reflects the connection relationship between multiple chords, which is the manifestation of the change of harmony with the performance, and can show the functional relationship and sound color between chords. Although an individual's emotional experience is affected by factors such as culture, background, and personal experience, some harmonic progressions will trigger common emotional reactions under normal circumstances. Therefore, in music works, the emotional ups and downs are often shown by designing the harmonic progression. For example, the [I - IV - V - I] harmonic progression is generally considered to have a stable and satisfying emotion and is commonly found in happy and celebratory music. The [i - iv - V - i] harmonic progression is generally considered to have a sad and melancholy emotion, especially when the chords contain minor keys, such as Am - Dm - E - Am.
[0050] After evaluating the emotional parameters of the target lyrics, the relevant information about the emotion that the user needs to express has been obtained, which can guide the emotional color of music creation. This step starts from the perspective of designing the harmonic progression to ensure that the generated music matches the user's emotional expression needs. Specifically, in order to establish the correlation between the emotional parameters and the harmonic progression, considering that there are relatively strict music theory rules in the production of existing music works, the emotions expressed by using the harmonic progression design can be accepted by the public. Therefore, a machine learning model can be used to learn the mapping relationship between the harmonic progression and the emotion parameters in existing music works. Through continuous training and parameter adjustment, the model can capture the mapping relationship between the emotional parameters and the harmonic progression that meets the emotional expression requirements. The model obtained thereby is the harmonic progression prediction model. The input of the harmonic progression prediction model is the emotion parameters, and the output is the target harmonic progression. The target harmonic progression is the harmonic progression predicted by the harmonic progression prediction model after learning the mapping relationship between the harmonic progression design of a large number of existing music works and the emotion parameters. The target harmonic progression matches the emotion parameters, and the emotion parameters are generated according to the target lyrics. Therefore, the obtained target harmonic progression will match the emotional expression needs of the target lyrics.
[0051] S106, input the target harmonic progression into the pitch prediction model to obtain the pitch probability sequence of each chord in the target harmonic progression.
[0052] It can be understood that in the process of automatically generating music, after generating a harmonic progression, it is necessary to select a suitable texture for the harmonic progression, and there may be problems such as out-of-tune and disharmony in the texture combinations generated by artificial intelligence. However, due to the lack of relevant criteria, it is difficult to determine whether the notes involved in the texture combination are qualified. For this reason, this embodiment needs to further confirm the pitch range of the target harmonic progression. Specifically, considering that existing music works will focus on note selection during production to avoid out-of-tune and disharmony problems, a machine learning model can be used to learn the preference relationship between each chord included in the harmonic progression and the corresponding pitch in existing music works. Through continuous training and parameter tuning, the model can capture the preference relationship between the harmonic progression and the selection of harmonious and in-tune notes. The model obtained thereby is the pitch prediction model. The input of the pitch prediction model is the target harmonic progression, and the output is the pitch probability sequence of each chord in the target harmonic progression. The pitch probability sequence corresponds one-to-one with each chord in the target harmonic progression, and it reflects the selection probability of each alternative note being selected as the note composition of the corresponding chord. For example, assuming that the samples learned by the pitch prediction model and the harmonic progression prediction model are unified in key, such as unified to C major, the alternative notes will only be the 12 notes corresponding to C major. At this time, the pitch probability sequence is the probability of these 12 notes being selected as the chord composition corresponding to them. In one embodiment, the input target harmonic progression is (F, G, Dm, Am), and the pitch probability sequence of the F chord is [0.072, 0.095, 0.081, 0.025, 0.123, 0.065, 0.042, 0.117, 0.036, 0.048, 0.119, 0.177].
[0053] S108. Determine the recommended notes and avoidance notes corresponding to each chord in the target harmonic progression according to the pitch probability sequence.
[0054] It can be understood that in this embodiment, the selection of the pitch range is refined into recommended notes and avoided notes. Recommended notes are the notes recommended for use. When there are recommended notes in the melody, no additional processing is required. Avoided notes are the notes that should not be used or should be used very rarely. Improper use of avoided notes will be considered disharmonious and out of tune. The pitch probability information can provide good suggestions for the final selection of melody notes. Among them, the notes with higher selection probabilities are the preferred choices in harmonious works, and the notes with lower selection probabilities are the choices to be avoided as much as possible in harmonious works. The recommended notes and avoided notes obtained based on the pitch probability information delimit the range for pitch selection, and can provide a criterion for judging whether the melody of the generated music is harmonious, so that timely processing can be carried out before the generated music is output to the user to generate a more reasonable and natural melody line. Regarding the use of recommended notes and avoided notes, it can be determined in combination with specific arrangement strategies. In some strategies, a corresponding texture combination is selected for the target harmonic progression. The avoided notes can be used to find out whether there are out-of-tune notes in the texture combination. If so, the out-of-tune notes are adjusted to one of the recommended notes. In some strategies, if there are avoided notes in the generated original melody, the avoided notes are masked, and then the masked melody is regenerated by the melody regeneration model to reduce the unreasonable use of avoided notes and generate a more harmonious and natural melody.
[0055] This solution first obtains the emotional parameters of the target lyrics input by the user in at least one emotional dimension. Then, the obtained emotional parameters are input into a specially trained harmonic progression prediction model to generate a target harmonic progression that matches the emotion of the target lyrics. The predicted target harmonic progression is input into another specially trained pitch prediction model to obtain the pitch probability sequence corresponding to each chord in the target harmonic progression. Finally, based on the pitch probability sequence, recommended notes and avoided notes are determined for each chord. This solution learns the music theory rules contained in existing music works through machine learning, effectively solving the problem that the generated music harmonic progression does not match the emotion that the user needs to express. At the same time, it provides a reference basis for solving the out-of-tune and disharmonious problems in artificial intelligence-generated music. Based on this solution, the works generated by artificial intelligence can meet higher quality requirements in terms of emotion expression, music rationality, etc., improving the feasibility and practicality of artificial intelligence music creation.
[0056] In one of the embodiments, the training process of the harmonic progression prediction model includes:
[0057] (1) Obtain a first training set with first annotation information. The first training set includes multiple first training harmonic progressions, and the first annotation information includes the emotional parameters of the first training harmonic progressions in at least one emotional dimension.
[0058] It can be understood that the first training harmonic progression is a harmonic progression extracted from existing musical works for training a harmonic progression prediction model. Machine learning models require a large amount of labeled data for training to learn the mapping relationship between the input and the output. The knowledge that the harmonic progression prediction model in this embodiment needs to learn is the mapping from emotional parameters to harmonic progressions. Therefore, it is necessary to obtain harmonic progression training data with emotional parameter annotations. For example, the selected emotional dimensions include pleasantness and arousal, and the value ranges of these two emotional dimensions are both 1-9. There is a first training harmonic progression (Am7, Fmaj7, Dm7, E7), and its corresponding first annotation information is (pleasantness: 1, arousal: 1). There is also a first training harmonic progression (C, G, Am, Em7), and its corresponding first annotation information is (pleasantness: 7, arousal: 7).
[0059] In addition, the number of first training harmonic progressions in each emotional classification quadrant of the first training set is greater than the preset lower limit of the number, and the difference in the number of first training harmonic progressions in different emotional classification quadrants is less than the deviation threshold.
[0060] (2) Unify the key of each first training harmonic progression in the first training set to the target key.
[0061] It can be understood that the key is the general term of the tonic of the key and the mode category. For example, C major is based on C as the tonic and major as the mode category. In this embodiment, in order to simplify the learning task of the model, the relative harmony theory is referred to as the theoretical framework. In this theoretical framework, chords are usually represented by their distance and relationship with the tonic chord. This method can make the structure of chord progressions easier to understand and can transplant chord progressions in different keys. That is, this theoretical framework pays more attention to the relative relationship of chord progressions rather than the specific key they are in. Therefore, for the same harmonic progression, they may eventually migrate to different keys, but their relative interval relationships are equivalent. Unifying the keys of all first training harmonic progressions to the target key will only reflect the relative relationship of chord progressions in all first training harmonic progressions, eliminating the additional training burden caused by different keys. In some embodiments, the target key is C major.
[0062] (3) Use the first training set with unified key to train the initial harmonic progression prediction model to obtain the harmonic progression prediction model.
[0063] It can be understood that the initial harmonic progression prediction model refers to the original harmonic progression prediction model using default parameters, and the first training set with unified tonality is the training material for the initial harmonic progression prediction model. The training process of this model is a typical supervised learning process. Taking the first annotation information as input, the predicted harmonic progression is obtained. The initial harmonic progression prediction model is adjusted with the goal of narrowing the difference between the predicted harmonic progression and the first training harmonic progression corresponding to the first annotation information, continuously improving the prediction performance of the model until the error requirement is finally met, and the final harmonic progression prediction model is obtained. The structure selected by the initial harmonic progression prediction model can be a transformer structure with 4 layers of encoders and 4 layers of decoders. This structure can achieve good results when performing the sequence-to-sequence (seq2seq) prediction task from emotion parameters to harmonic progressions.
[0064] In one embodiment, obtaining the first training set with the first annotation information includes:
[0065] (1) Obtain the first training harmonic progression and its corresponding initial emotion parameter set. The initial emotion parameter set includes the initial emotion parameters annotated for the first training harmonic progression by multiple annotation subjects in at least one emotion dimension.
[0066] It can be understood that the collection of the first training set requires two parts, namely the collection of training samples and the annotation of training samples. Among them, the first training harmonic progression can be extracted from the MIDI file of the piano accompaniment work. After obtaining the first training harmonic progression, it is necessary to evaluate the first training harmonic progression from at least one selected emotion dimension to obtain the corresponding emotion parameters. Since the evaluation of emotions is highly subjective, in order to ensure the generality of the annotation information, in this embodiment, multiple annotation subjects are used to annotate each first training harmonic progression together. The annotation results obtained by these annotation subjects are the initial emotion parameters, and the set formed by the initial emotion parameters is the initial emotion parameter set. The initial emotion parameter set needs to go through further comprehensive analysis to obtain the final first annotation information. In addition, the annotation subjects here can be annotators or annotation models. That is, the initial emotion parameter set can be manually annotated, semi-automatically annotated, or fully automatically annotated.
[0067] (2) For any emotion dimension, screen out the invalid annotations in the initial emotion parameters, and obtain the emotion parameter corresponding to this emotion dimension according to the first average value of the initial emotion parameters after screening.
[0068] It can be understood that since the initial emotion parameter set may include initial emotion parameters corresponding to more than one emotion dimension, each emotion dimension needs to be screened separately. The screening is to exclude the initial emotion parameters that deviate from other annotation subjects beyond the set conditions and retain the initial emotion parameters that can reach a consensus, so that the finally obtained first annotation information is relatively objective and more in line with the preferences of the public. The first average obtained after averaging the screened initial emotion parameters is the emotion parameter corresponding to the emotion dimension. Among them, the determination of invalid annotations can be to first calculate the average of all initial emotion parameters under each emotion dimension, that is, to obtain the second average, and also calculate the standard deviation of all initial emotion parameters. Then, each initial emotion parameter to be screened is compared with the second average respectively. If the absolute value of the difference from the second average is greater than the standard deviation, it can be determined that the degree of deviation of the initial emotion parameter from the annotation results of other annotation subjects is relatively large and does not meet the requirements. The initial emotion parameter can be determined as an invalid annotation.
[0069] In one of the embodiments, the training process of the pitch prediction model includes:
[0070] (1) Obtain a second training set with second annotation information. The second training set includes a plurality of second training harmonic progressions with the tonality unified to the target tonality, and the second annotation information includes the notes contained in each chord in the second training harmonic progression.
[0071] It can be understood that the second training harmonic progression is a harmonic progression extracted from existing music works for training the pitch prediction model. The machine learning model needs a large amount of annotated data for training to learn the mapping relationship between the input and the output. The knowledge that the pitch prediction model in this embodiment needs to learn is the mapping between the harmonic progression and the preference selection of chord notes. Therefore, it is necessary to obtain harmonic progression training data containing note selection annotations. For example, if the second training harmonic progression is (Am, C, Am, G), the corresponding second annotation information needs to reflect the notes contained in each chord, and the second annotation information can be represented in a vectorized form. In the vector, each bit corresponds to one of the 12 alternative notes respectively. If the corresponding bit is 1, it means that the chord contains the note at the corresponding position. If the corresponding bit is 0, it means that the chord does not contain the note at the corresponding position. For example, Am: [0, 0, 1, 0, 1, 0, 1, 0, 0, 0, 1, 0]. Since the harmonic progression prediction model focuses more on the relative relationship of the chord progression rather than the specific tonality in the relative harmony theory framework, in order to be consistent with the previous model and reduce the learning difficulty, the second training harmonic progression in this embodiment also needs to unify the tonality to the target tonality. The second training harmonic progression here can be extracted from a piece of music different from the first training harmonic progression, or it can be the same.
[0072] (2)Train the initial pitch prediction model using the second training set to obtain the pitch prediction model.
[0073] It can be understood that the initial pitch prediction model refers to the original pitch prediction model using default parameters, and the second training set is the training material for the initial pitch prediction model. Different from the sequence-to-sequence prediction model of the harmonic progression prediction model, the initial pitch prediction model is a multi-classification prediction model. Its input is the harmonic progression, and the output is the predicted pitch probability sequence corresponding to the harmonic progression. The training of this type of model generally calculates the cross-entropy loss function based on the second annotation information and the predicted pitch probability sequence, and adjusts the parameters of the initial pitch prediction model until the value of the cross-entropy loss function converges or meets the requirements to obtain the final pitch prediction model. The structure selected for the initial pitch prediction model can be a two-layer bilstm model.
[0074] In one embodiment, according to the pitch probability sequence, determining the recommended notes corresponding to each chord in the target harmonic progression includes:
[0075] (1)Determine the in-chord notes and out-of-chord notes according to the corresponding chord types.
[0076] It can be understood that the in-chord notes refer to the constituent notes of the chord, which are determined by the chord type of the chord. The out-of-chord notes refer to other notes in the target key except the in-chord notes. For example, in the Cmaj chord in C major, the three notes C, E, and G are the in-chord notes. The notes other than C, E, and G among the twelve notes corresponding to C major are the out-of-chord notes.
[0077] (2)Take the in-chord notes among the top three notes with the highest probabilities in the pitch probability sequence and the out-of-chord notes among the top five notes with the highest probabilities as the recommended notes.
[0078] It can be understood that the pitch probability sequence reflects the frequency of alternative notes appearing as the notes of this chord in the training set. A high probability means a high chance of being accepted and used in a large amount of music. As the constituent notes of the chord, the in-chord notes should be used most frequently in the arrangement. Therefore, take the in-chord notes among the top three notes with the highest probabilities as the recommended notes. And the out-of-chord notes can be further divided into notes that can be used to increase the complexity and expressiveness of the chord, and notes that damage the harmony. The former type of out-of-chord notes can make the chord more rich and is recommended to be used in the chord. While the latter type of out-of-chord notes should be avoided in the chord. The former type of out-of-chord notes will also be used more frequently in this chord, and this frequent use will also be reflected in the pitch probability sequence, making it have a greater selection probability. Therefore, take the out-of-chord notes among the top five notes with the highest probabilities as the recommended notes.
[0079] In one embodiment, determining the avoidance tone corresponding to each chord in the target harmony progression according to the pitch probability sequence includes:
[0080] (1) Determine the out-of-tune notes based on the target tonality.
[0081] It can be understood that each target tonality will correspond to twelve notes, and the notes outside these twelve notes are out-of-tune notes. Out-of-tune notes themselves are dissonant parts in the structure of the music, and their use needs to have a clear purpose and meaning.
[0082] (2) Determine the chord tones based on the corresponding chord type.
[0083] It can be understood that the chord tones refer to the constituent tones of the chord, which are determined by the chord type of the chord. The chord tones outside the chord refer to the other notes in the target tonality except the chord tones. For example, in the Cmaj chord in C major, the three notes C, E, and G are the chord tones. The notes other than C, E, and G in the twelve notes corresponding to C major are the chord tones outside the chord.
[0084] (3) The out-of-key tones and the out-of-chord tones in the first two notes with the smallest probability among the notes other than the out-of-key tones in the probability sequence of pitch are taken as avoidance tones.
[0085] It can be understood that the pitch probability sequence reflects the frequency of candidate notes appearing as notes of the chord in the training set. A high probability means a high probability of being accepted and used in a large amount of music. The chord tones can be further divided into notes that can be used to increase the complexity and expressiveness of the chord, and notes that destroy the harmony. The former chord tones can make the chord richer and are recommended for use in the chord. The latter chord tones should be avoided in the chord. The latter chord tones will rarely appear in the chord, and this cautious use will also be reflected in the pitch probability sequence, making it have a smaller selection probability. Therefore, the chord tones in the smallest first two notes are used as avoidance tones. In addition, if there is a conflict between the avoidance tone and the selection and recommendation of the tone, the note is preferentially determined as the avoidance tone.
[0086] The present application embodiment provides a chord prediction device, see Figure 2 , including a data acquisition module 210, a harmony progression prediction module 220, a pitch probability prediction module 230 and a pitch range determination module 240.
[0087] The data acquisition module 210 is used to acquire the emotional parameters of the target lyrics in at least one emotional dimension.
[0088] The harmony progression prediction module 220 is used to input the emotion parameters into the harmony progression prediction model to obtain the target harmony progression.
[0089] The pitch probability prediction module 230 is used to input the target harmony progression into the pitch prediction model to obtain the pitch probability sequence of each chord in the target harmony progression.
[0090] The pitch range determination module 240 is used to determine the recommended pitches and avoidance pitches corresponding to each chord in the target harmony progression according to the pitch probability sequence.
[0091] In one embodiment, the chord prediction device further includes a first model training module. The first model training module is used to obtain a first training set with first annotation information; the first training set includes a plurality of first training harmony progressions, and the first annotation information includes the emotional parameters of the first training harmony progressions in at least one emotional dimension; unify the key of each first training harmony progression in the first training set to the target key; use the first training set with unified key to train the initial harmony progression prediction model to obtain the harmony progression prediction model.
[0092] In one embodiment, the first model training module is used to obtain the first training harmony progression and its corresponding initial emotional parameter set; the initial emotional parameter set includes the initial emotional parameters annotated by a plurality of annotation subjects for the first training harmony progression in at least one emotional dimension; for any emotional dimension, filter out the invalid annotations in the initial emotional parameters, and obtain the emotional parameter corresponding to this emotional dimension according to the first average value of the filtered initial emotional parameters.
[0093] In one embodiment, the first model training module is used to determine the second average value and standard deviation of all initial emotional parameters in any emotional dimension; if the absolute value of the difference between the initial emotional parameter and the second average value is greater than the standard deviation, then determine this initial emotional parameter as an invalid annotation.
[0094] In one embodiment, the chord prediction device further includes a second model training module. The second model training module is used to obtain a second training set with second annotation information; the second training set includes a plurality of second training harmony progressions whose keys are unified to the target key, and the second annotation information includes the notes contained in each chord in the second training harmony progression; use the second training set to train the initial pitch prediction model to obtain the pitch prediction model.
[0095] In one embodiment, the pitch range determination module 240 is used to determine the in-chord notes and out-of-chord notes according to the corresponding chord type; arrange them in descending order of probability, and select the in-chord notes among the top three notes with the highest probability in the pitch probability sequence and the out-of-chord notes among the top five notes with the highest probability as the recommended pitches.
[0096] In one embodiment, the pitch range determination module 240 is configured to determine out-of-key tones according to the target key; determine non-chord tones according to the corresponding chord types; and use the out-of-key tones and the non-chord tones among the two notes with the lowest probabilities in the pitch probability sequence except for the out-of-key tones as avoidance tones.
[0097] For the specific limitations on the chord prediction device, reference can be made to the limitations on the chord prediction method in the above text, which will not be elaborated here. Each module in the above chord prediction device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0098] The embodiments of the present application provide a computer device, including one or more processors and a memory. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by one or more processors, the steps of the chord prediction method in any of the above embodiments are executed.
[0099] Schematically, as Figure 3 shown, Figure 3 is an internal structure schematic diagram of a computer device provided by an embodiment of the present application. Referring to Figure 3 , the computer device 300 includes a processing component 302, which further includes one or more processors, and memory resources represented by a memory 301 for storing instructions executable by the processing component 302, such as application programs. The application programs stored in the memory 301 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 302 is configured to execute instructions to perform the steps of the chord prediction method in any of the above embodiments.
[0100] The computer device 300 may further include a power supply component 303 configured to perform power management of the computer device 300, a wired or wireless model interface 304 configured to connect the computer device 300 to a model, and an input / output (I / O) interface 305.
[0101] The embodiments of the present application provide a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the chord prediction method in any of the above embodiments.
[0102] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0103] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.
[0104] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A chord prediction method, characterized in that: include: Obtaining emotional parameters of target lyrics in at least one emotional dimension; Inputting the emotion parameter into a harmony progression prediction model to obtain a target harmony progression; Inputting the target harmony progression into a pitch prediction model to obtain a pitch probability sequence of each chord in the target harmony progression; According to the pitch probability sequence, the recommended tones and the avoided tones corresponding to the chords in the target harmony progression are determined.
2. The chord prediction method according to claim 1, characterized in that: The training process of the harmony progression prediction model includes: Acquire a first training set with first annotation information; the first training set includes a plurality of first training harmony progressions, and the first annotation information includes the emotion parameters of the first training harmony progressions in the at least one emotion dimension; unifying the tonality of each of the first training harmony progressions in the first training set into a target tonality; The initial harmony progression prediction model is trained using the first training set with unified tonality to obtain the harmony progression prediction model.
3. The chord prediction method according to claim 2, characterized in that: The step of obtaining a first training set with first annotation information includes: Acquire the first training harmony progression and its corresponding initial emotion parameter set; the initial emotion parameter set includes initial emotion parameters annotated by multiple annotation subjects on the first training harmony progression in at least one emotion dimension; For any of the emotion dimensions, invalid annotations in the initial emotion parameters are screened out, and the emotion parameter corresponding to the emotion dimension is obtained based on a first average of the screened initial emotion parameters.
4. The chord prediction method according to claim 3, characterized in that: For any of the emotion dimensions, filtering out invalid annotations in the initial emotion parameters includes: For any of the emotion dimensions, determining a second mean and a standard deviation of all the initial emotion parameters under the emotion dimension; If the absolute value of the difference between the initial emotion parameter and the second average is greater than the standard deviation, the initial emotion parameter is determined as the invalid label.
5. The chord prediction method according to claim 2, characterized in that: The number of the first training harmony progressions in each emotion classification quadrant of the first training set is greater than a preset lower limit, and the difference in the number of the first training harmony progressions in different emotion classification quadrants is less than a deviation threshold.
6. The chord prediction method according to claim 2, characterized in that: The training process of the pitch prediction model includes: Acquire a second training set with second annotation information; the second training set includes a plurality of second training harmonic progressions whose tonality is unified into the target tonality, and the second annotation information includes notes included in each chord in the second training harmonic progression; The initial pitch prediction model is trained using the second training set to obtain the pitch prediction model.
7. The chord prediction method according to claim 1, characterized in that: Determining the recommended tones corresponding to the chords in the target harmony progression according to the pitch probability sequence includes: Determine the tones within and outside the chord according to the corresponding chord types; The chord tones among the first three notes with the highest probability in the pitch probability sequence and the chord tones outside the chord among the first five notes with the highest probability in the pitch probability sequence are used as the recommended tones.
8. The chord prediction method according to claim 5, characterized in that: Determining the avoidance tones corresponding to the chords in the target harmony progression according to the pitch probability sequence includes: determining an out-of-key tone according to the target tonality; Determine the out-of-chord tone according to the corresponding chord type; The out-of-tune tone and the out-of-chord tone in the first two notes with the smallest probability among the notes other than the out-of-tune tone in the pitch probability sequence are used as the avoidance tone.
9. A chord prediction device, characterized in that: include: A data acquisition module, used to acquire the emotional parameters of the target lyrics in at least one emotional dimension; A harmony progression prediction module, used for inputting the emotion parameter into a harmony progression prediction model to obtain a target harmony progression; A pitch probability prediction module, used for inputting the target harmony progression into a pitch prediction model to obtain a pitch probability sequence of each chord in the target harmony progression; The pitch range determination module is used to determine the recommended tones and avoidance tones corresponding to each chord in the target harmony progression according to the pitch probability sequence.
10. A computer device, characterized in that: The method comprises one or more processors and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the chord prediction method according to any one of claims 1 to 8 are executed.
11. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the chord prediction method according to any one of claims 1 to 8.