A method for identifying musical instrument music signals and chords

Through audio feature extraction and Fourier transform, combined with the filter group of Mel frequency scale, the audio signals of the instrument are processed, a two-dimensional matrix representation is constructed and a feature extraction model is input. Finally, the chord function recognition model is used to identify chord features, which solves the problem of chord recognition relies on manual and low accuracy in the existing technology, and realizes automatic chord recognition and improves recognition accuracy.

CN115083373BActive Publication Date: 2025-05-30GUANGZHOU RANTION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210594579.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-05-30
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

In the prior art, chords in music scores need to be manually identified and annotated, and the accuracy of chord recognition is low.

Method used

The audio analog signal of the instrument is converted into a digital signal through the radio device, and the audio feature module is used to extract the unmanned voice features, and the frequency domain signal is segmented through the filter group of Fourier transform and Mel frequency scale to obtain note information. Then, a two-dimensional matrix representation sequence is constructed and a feature extraction model is input to obtain the note feature sequence. Finally, the note feature sequence is input to the chord function recognition model, the chord features are identified through chord mode, tone, and index dimensions, and the model parameters are adjusted through prediction loss.

Benefits of technology

Automatic recognition of chords of music scores is realized, the accuracy of chord recognition is improved, and the dependence on manual recognition is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115083373B_ABST
    Figure CN115083373B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying musical instrument music signals and chords, comprising the following steps: signal conversion, constructing a note feature sequence, comparing chord features, and chord correction. The audio features are converted into unvoiced features through an audio feature module, and then the Fourier transform is used to segment the unvoiced features to obtain note information. Then, chord features of each note in different chord function recognition dimensions are extracted through a note feature model, and the chord features are continuously recognized through the note feature model, so as to realize the automatic recognition of musical score chords. Moreover, the training data set, the test data set, and the validation data set will continuously train and optimize the chord features, thereby improving the accuracy of chord recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of audio and video, and particularly relates to a method for identifying musical instrument music signals and chords. Background Art

[0002] Automatic chord recognition (ACR) is an important research topic in the field of music information retrieval. Chords are the basis for the formation of music. They can directly affect the aesthetics of a piece of music, determine the emotion and color of the music, and enrich the content of the music. Common chord types include triads: major triad (maj), minor triad (min), diminished triad (dim), augmented triad (aug), seventh chords: major seventh chord (maj7), dominant seventh chord (7), minor seventh chord (m7), minor major seventh chord (mM7), diminished seventh chord (dim7), augmented seventh chord (aug7), half-diminished seventh chord (half-dim7), as well as added chords (add), extended chords (extend), and suspended chords (suspend);

[0003] Chord progressions are the basis for popular music, classical music, rock music, blues, jazz, and other types of music. In these music genres, melodies and rhythms are built on chord progressions. During the songwriting process, a song usually starts with a chord progression, and then other elements such as melody, bass, and instrumentation are considered. Therefore, chord recognition is undoubtedly very important for music learning and composition. For example, in jazz improvisation, a jazz bassist can improvise the bass according to the chord sheet, and other solo instrument players can also improvise their own solo parts according to the chord sheet. In the related art, some websites provide some manually annotated sheet music for download. However, the chords in these sheet music need to be identified and annotated manually, and the accuracy of chord recognition is relatively low. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for identifying musical instrument music signals and chords to solve the problems in the prior art that the chords in the sheet music need to be identified and annotated manually and the accuracy of chord recognition is relatively low as mentioned in the above background art.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A method for identifying musical instrument music signals and chords, comprising the following steps:

[0006] S1. Convert the audio analog signal of the musical instrument into a digital signal through a radio receiving device, transmit it to an audio feature module, extract the audio features in the digital signal through the audio feature module to obtain a non-vocal feature, then perform a Fourier transform on the time-domain signal of the non-vocal feature to convert it to the frequency domain, and then use a filter bank with Mel frequency scale to segment the frequency-domain signal so that each frequency band corresponds to a value;

[0007] S2. Then, extract features from the values after splitting the frequency-domain signal to obtain note information, construct a two-dimensional matrix representation for each note from the note information, and then input the two-dimensional matrix representation sequence into the feature extraction model to obtain the note feature sequence output by the feature extraction model for the two-dimensional matrix representation sequence. The note feature sequence contains the note features corresponding to each note.

[0008] S3. Then, input the note feature sequence into the chord function recognition model, and establish a note feature model to obtain the chord features obtained by recognizing and processing each note feature in the note feature sequence from different chord function recognition dimensions.

[0009] S4. Finally, based on the true chord type of the original audio, calculate the prediction loss of the note feature model, and then adjust the parameters of the note feature model based on the prediction loss of the note feature model.

[0010] Preferably, in S1, the correspondence between the Mel frequency f mel and the sound signal frequency f is:

[0011] F mel = 2959log 10 (1 + f / 700).

[0012] Preferably, in S2, the feature extraction model converts audio data from the audio analog signal and compares and identifies it with the note features. The note features include at least one set of corresponding relationships between note segments and feature information.

[0013] Preferably, the audio data is trained by the feature extraction model to obtain sheet music information, the notes in the sheet music information are compared and identified with the note features, and then the note features with correct comparison results are saved, and the note features with incorrect comparison results are screened out.

[0014] Preferably, in S3, the chord function recognition dimensions include the chord mode dimension, the chord tonality dimension, and the chord inversion dimension. The chord mode dimension, the chord tonality dimension, and the chord inversion dimension jointly act on the chord function representation of the original audio analog signal.

[0015] Preferably, the note feature model in S3 contains the correct note features from S2. A note database is established in the note feature model to save the correct note features, and the note feature model will continuously train and optimize the note features in the note database.

[0016] Preferably, in S4, the audio feature model divides the correct note features in the note database into a first note data segment, a second note data segment, and a third note data segment. The first note data segment constitutes the training data set, the second note data segment constitutes the test data set, and the third note data segment constitutes the validation data set. Then, the training data set, the test data set, and the validation data set are used to predict the note features.

[0017] Preferably, the training data set trains the note features, tests them through the test data set, and summarizes the verification results after re-verification through the validation data set to obtain the chord recognition result.

[0018] The technical effects and advantages of the present invention: A method for identifying musical instrument music signals and chords proposed by the present invention has the following advantages compared with the prior art:

[0019] The audio feature module converts the audio features into non-vocal features, then uses Fourier transform to segment the non-vocal features to obtain note information, and then extracts the chord features of each note in different chord function recognition dimensions through the note feature model. The chord features are continuously recognized through the note feature model, so as to realize the automatic recognition of musical score chords. Moreover, the training data set, the test data set, and the validation data set continuously train and optimize the chord features, thereby improving the accuracy of chord recognition. Description of the Drawings

[0020] Figure 1 It is a relationship diagram of the Mel frequency and the standard frequency of the present invention. Detailed Embodiments

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The specific embodiments described here are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] The present invention provides a method for identifying musical instrument music signals and chords as Figure 1 shown, including the following steps:

[0023] S1. Signal conversion: The audio analog signal of the musical instrument is converted into a digital signal through a sound collection device and transmitted to the audio feature module. The audio features in the digital signal are extracted by the audio feature module to obtain the non-vocal features. Then, the time-domain signal of the non-vocal features is subjected to Fourier transform to be converted into the frequency domain. Then, the filter bank with Mel frequency scale is used to segment the corresponding frequency-domain signal so that each frequency band corresponds to a value. The correspondence between the Mel frequency fmel and the sound signal frequency f is as follows:

[0024] F mel = 2959 log 10 (1 + f / 700);

[0025] If it is evenly divided on the Mel scale, the distance between the corresponding Herz will become larger and larger. The meaning of the cepstrum is: perform Fourier transform on the time-domain signal, then take the log, and then perform the inverse Fourier transform, which can be divided into complex cepstrum, real cepstrum, and power cepstrum;

[0026] S2. Construct the note feature sequence: Then, extract features from the values after splitting the frequency-domain signal to obtain note information. Construct a two-dimensional matrix representation for each note from the note information, and then input the sequence of two-dimensional matrix representations into the feature extraction model to obtain the note feature sequence output by the feature extraction model for the sequence of two-dimensional matrix representations. The note feature sequence contains the note features corresponding to each note. The feature extraction model converts audio data from the audio analog signal and compares and identifies it with the note features. The note features include at least one set of corresponding relationships between note segments and feature information. The audio data is trained by the feature extraction model to obtain sheet music information. Compare and identify the notes in the sheet music information with the note features, then save the note features with correct comparison results and filter out the note features with incorrect comparison results. If the note pitch and note duration of each note are extracted for each note, the note pitch can be used as the vertical element in the two-dimensional matrix, and the note duration can be used as the horizontal element in the two-dimensional matrix, or the note pitch can be used as the horizontal element in the two-dimensional matrix, and the note duration can be used as the vertical element in the two-dimensional matrix to construct the two-dimensional matrix representation of each note. For the two-dimensional matrix representation constructed by using the note pitch as the vertical element and the note duration as the horizontal element in the two-dimensional matrix, it can be understood as a coordinate system constructed with the note pitch as the vertical coordinate and time as the horizontal coordinate. Therefore, the two-dimensional matrix representation of each note can carry the note information corresponding to the note, and further obtain the relevant information of the music data in terms of chord function based on the recognition and processing of the two-dimensional matrix representation of each note. The feature extraction model includes a Bi-LSTM network and a fully connected network. The Bi-LSTM network is used to extract note features from the sequence of two-dimensional matrix representations input into it. Among them, the sequence of two-dimensional matrix representations is a sequence composed of the two-dimensional matrix representations of each note contained in the music data of the music chord to be recognized, and the two-dimensional matrix representation of each note is constructed based on the note information corresponding to each note;

[0027] S3. Compare the chord features: Then, input the note feature sequence into the chord function recognition model and establish a note feature model to obtain the chord features obtained by recognizing and processing each note feature in the note feature sequence from different chord function recognition dimensions. The chord function recognition dimensions include the chord mode dimension, the chord tonality dimension, and the chord inversion dimension. The chord mode dimension, the chord tonality dimension, and the chord inversion dimension jointly act on the chord function representation of the original audio analog signal. The note feature model contains the correct note features from S2. A note database is established in the note feature model to save the correct note features, and the note feature model will continuously train and optimize the note features in the note database;

[0028] S4, Chord Correction: Finally, based on the true chord type of the original audio, calculate the prediction loss of the note feature model, and then adjust the parameters of the note feature model based on the prediction loss of the note feature model. The audio feature model divides the correct note features in the note database into the first note data segment, the second note data segment, and the third note data segment. The first note data segment constitutes the training data set, the second note data segment constitutes the test data set, and the third note data segment constitutes the validation data set. Then, use the training data set, the test data set, and the validation data set to predict the note features. The training data set will train the note features, and then test them through the test data set. After passing the validation again through the validation data set, summarize the validation results to obtain the chord recognition result. By using the training data set, the test data set, and the validation data set to predict the note features, the accuracy of chord recognition is improved.

[0029] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for identifying musical instrument music signals and chords, comprising the following steps: Characterized in that: S1. Signal conversion: The audio analog signal of the musical instrument is converted into a digital signal through a sound collection device and transmitted to the audio feature module. The audio features in the digital signal are extracted by the audio feature module to obtain the non-vocal features. Then, the time-domain signal of the non-vocal features is subjected to Fourier transform to be converted into the frequency domain. Then, a filter bank with Mel frequency scale is used to segment the frequency-domain signal so that each frequency segment corresponds to a value; S2. Constructing a note feature sequence: Then, the values after segmentation of the frequency-domain signal are subjected to feature extraction to obtain note information. The note information is used to construct a two-dimensional matrix representation of each note. Then, the two-dimensional matrix representation sequence is input into the feature extraction model to obtain the note feature sequence output by the feature extraction model for the two-dimensional matrix representation sequence. The note feature sequence contains the note features corresponding to each note; Among them, constructing a two-dimensional matrix of each note from the note information specifically means: using the note pitch as the vertical element in the two-dimensional matrix and the note duration as the horizontal element in the two-dimensional matrix, or using the note pitch as the horizontal element in the two-dimensional matrix and the note duration as the vertical element in the two-dimensional matrix to construct the two-dimensional matrix representation of each note; S3. Comparing chord features: Then, the note feature sequence is input into the chord function recognition model, and a note feature model is established to obtain the chord features obtained by recognizing and processing each note feature in the note feature sequence from different chord function recognition dimensions; S4. Chord correction: Finally, based on the true chord type of the original audio, the prediction loss of the note feature model is calculated, and then the parameters of the note feature model are adjusted based on the prediction loss of the note feature model.

2. A method for identifying musical instrument music signals and chords according to claim 1, Characterized in that: The corresponding relationship between the Mel frequency f in S1 mel and the frequency f of the sound signal is as follows: F mel = 2959 log 10 (1 + f / 700).

3. A method for identifying musical instrument music signals and chords according to claim 1, Characterized in that: In S2, the feature extraction model converts audio data from the audio analog signal and compares and identifies it with the note features. The note features include at least one set of corresponding relationships between note segments and feature information.

4. A method for identifying musical instrument music signals and chords according to claim 3, Characterized in that: The audio data is trained by the feature extraction model to obtain score information. The notes in the score information are compared and identified with the note features, and then the note features with correct comparison results are saved, and the note features with incorrect comparison results are screened out.

5. A method for identifying musical instrument music signals and chords according to claim 1, Characterized in that: The chord function recognition dimensions in S3 include the chord mode dimension, the chord tonality dimension, and the chord inversion dimension. The chord mode dimension, the chord tonality dimension, and the chord inversion dimension jointly act on the chord function representation of the original audio analog signal.

6. A method for identifying musical instrument music signals and chords according to claim 1, Characterized in that: The correct note features from S2 are included in the note feature model in S3. A note database is established in the note feature model to save the correct note features, and the note feature model continuously trains and optimizes the note features in the note database.

7. A method for identifying musical instrument music signals and chords according to claim 6, wherein: In S4, the audio feature model divides the correct note features in the note database into a first note data segment, a second note data segment, and a third note data segment. The first note data segment constitutes a training data set, the second note data segment constitutes a test data set, and the third note data segment constitutes a validation data set, and the training data set, the test data set, and the validation data set are used to predict the note features.

8. A method for identifying musical instrument music signals and chords according to claim 7, wherein: The training data set trains the note features, tests them through the test data set, and after being verified again through the validation data set, the verification results are summarized to obtain the chord recognition result.

Citation Information

Patent Citations

  • Chord recognition method combining SVM with enhanced PCP

    CN103714806A

  • Advertisement putting system and method for cloud electronic commerce

    CN111582951A