Method and device for judging western tones and five tones
By pre-processing and leading tone determination, combined with specific processing of Western size and national five-tone modes, the problems of high complexity and small recognition range in the existing technology are solved, and more accurate and efficient mode judgment is achieved.
Patent Information
- Application Number
- CN202510133709.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The prior art identification method between the five-tone tone style, Western size tone and other modes has high complexity, a small recognition range, and cannot effectively identify other modes except the five-tone mode.
By obtaining the target audio and preprocessing, extracting audio information, determining the main voice, and confirming the mode category tag based on the main voice. For Western large and small modes and national five-tone modes, we adopt corresponding strategies to make judgments, and fully consider the uniqueness of different mode systems.
It improves the accuracy and adaptability of mode judgment, reduces the judgment workload, simplifies data processing, reduces the computational complexity, and enhances the recognition ability of different modes.
Smart Images

Figure CN119964529A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer music information processing, and in particular to a method and device for judging western major and minor keys and pentatonic modes. Background Art
[0002] In the field of music mode recognition, one type relies entirely on manual processing to judge the tune, and the other type relies entirely on machine learning models to judge the tune based on the rapidly developing machine learning technology. In the related technologies, the statistical market of each note is used to judge the pattern, or the national mode judgment is based entirely on music knowledge for national mode, and only the most basic processing is performed on the pentatonic mode, and then the machine learning is used to generate the model for recognition.
[0003] However, in the related art, the pattern judgment is completely based on the statistical market of each note. Since there may be similar situations in different modes only from the perspective of note duration, the judgment is likely to be wrong; for ethnic modes, the ethnic mode judgment based entirely on music knowledge is used, which is only effective for pentatonic modes and cannot be used to identify other modes, resulting in a large amount of waste of early processing information; only the most basic processing is performed on pentatonic modes, and the method of model recognition is handed over to machine learning. Only in pentatonic modes, this machine learning can be completely replaced by the summarized music knowledge, and there is no inevitable advantage. In summary, the existing technology in the recognition between pentatonic modes and Western major and minor modes and other modes either focuses entirely on a certain field or considers all modes in general, resulting in a small range of recognized music modes and high complexity of the recognition method, which urgently needs to be improved. Summary of the invention
[0004] The present application provides a method and device for judging Western major and minor keys and pentatonic modes, so as to solve the problem that the related art either focuses entirely on a certain field or considers all modes in general in the recognition between pentatonic modes and Western major and minor keys and other modes, resulting in a small range of recognized music modes and high complexity of the recognition method.
[0005] The first aspect of the present application provides a method for judging Western major or minor key and pentatonic mode, comprising the following steps: obtaining a target audio to be judged, and preprocessing the target audio to obtain audio information of the target audio; determining the tonic of the target audio according to the audio information, and confirming the mode category label of the target audio according to the tonic; confirming the mode of the target audio according to the mode category label, and when the mode is a Western major or minor mode or a national pentatonic mode, adopting a corresponding strategy to obtain a mode judgment result.
[0006] Through the above technical solution, the embodiment of the present application can obtain the target audio and perform preprocessing to obtain audio information, remove noise and other interference, accurately extract effective music features, and provide a high-quality data basis for subsequent accurate judgment. Comprehensively consider multiple music elements to determine the tonic based on audio information, so that the tonic judgment is more accurate and reliable; then confirm the mode category label based on the tonic, effectively narrow the mode range, and reduce the judgment workload; finally, adopt corresponding strategies for Western major and minor modes and national pentatonic modes to obtain judgment results, fully consider the uniqueness of different mode systems, and improve the accuracy and adaptability of judgment.
[0007] Optionally, in one embodiment of the present application, the target audio is preprocessed to obtain audio information of the target audio, including: extracting basic beat information in the music based on the target audio; separating the tracks in the target audio, determining the main instrument track according to the duration proportion of each track in the target audio and the uniqueness of each track, and extracting the duration information of all pitches based on the main instrument track; retaining a single melody and removing duplications based on the duration information to obtain a non-repetitive sequence in the main track.
[0008] Through the above technical scheme, the embodiment of the present application can provide key rhythmic basis for subsequent processing by extracting basic beat information of the music, so that the entire music processing process is more in line with the rhythmic characteristics of the music itself, and the music theory basis of the scheme is enhanced. The main instrument track is determined according to the total duration ratio and uniqueness of each separated track, which can accurately locate the source of the core melody, avoid interference from other tracks, and ensure that subsequent analysis focuses on key music elements. Retaining a single melody and removing duplicate operations based on duration information simplifies the data and reduces the impact of redundant information on judgment. It not only reduces the computational complexity and improves processing efficiency, but also makes the extracted non-repetitive sequence more accurately represent the essential melodic characteristics of the music, laying a solid and reliable data foundation for subsequent steps such as accurately determining the tonic and judging the mode, and improving the accuracy and effectiveness of the entire mode judgment technical solution.
[0009] Optionally, in one embodiment of the present application, determining the tonic of the target audio according to the audio information includes: calculating in sequence the note frequencies of the same piece of music under different rhythmic patterns according to the beat distribution of different beat types; and determining the tonic based on the note frequency statistics of the highest frequency notes and their frequencies after adjusting to standardized twelve tones.
[0010] Through the above technical scheme, the embodiment of the present application can calculate the note frequency by distributing the beats according to different beat types, so that the calculation result is more in line with the actual music performance and auditory perception. The tonic is determined based on the highest frequency note and its frequency after the statistical standardization of the twelve tones, which comprehensively integrates the rhythm and pitch frequency information and avoids the limitation of judging the tonic by a single factor. When dealing with music with complex rhythm changes or irregular melody trends, the tonic can be locked more accurately, providing a solid and reliable foundation for subsequent mode judgment, and improving the accuracy and stability of the entire mode judgment process.
[0011] Optionally, in one embodiment of the present application, when the mode is a Western major or minor mode or a national pentatonic mode, a corresponding strategy is adopted to obtain a mode judgment result, including: when the mode is the Western major or minor mode, counting all the first pitches in the main track; taking the tonic as a reference, standardizing all the first pitches to a sequence of 12 semitones above the tonic, and counting the first frequency of the duration of occurrence of each pitch; confirming the major or minor key, key signature and tonality according to the first frequency of the duration of occurrence of each pitch to obtain the mode judgment result.
[0012] Through the above technical scheme, the embodiment of the present application can count all the first pitches in the main track and standardize them, unify the pitch data in a standard framework that is easy to analyze, and make the subsequent frequency-based analysis more consistent and comparable. According to the first frequency of the duration of each pitch, key elements such as major and minor keys, key signatures and tonality are confirmed, and the pitch duration frequency information is fully utilized. This quantitative analysis method is more scientific and accurate than traditional qualitative or empirical judgment. It can capture the distribution characteristics of different pitches in the time dimension in music more meticulously, so as to accurately determine the relevant attributes of the mode, whether it is to distinguish the subtle differences between major and minor keys, or to determine the key signature and judge the tonality. It is more reliable, which greatly improves the accuracy and efficiency of Western major and minor mode judgment.
[0013] Optionally, in one embodiment of the present application, when the mode is a Western major or minor mode or a national pentatonic mode, a corresponding strategy is adopted to obtain a mode judgment result, including: when the mode is the national pentatonic mode, counting all second pitches in the main track; taking the tonic as a reference, standardizing all second pitches to a sequence of 12 semitones above the tonic, and counting the second frequency of the duration of occurrence of each pitch; confirming the mode category and mode according to the second frequency of the duration of occurrence of each pitch to obtain the mode judgment result.
[0014] Through the above technical solution, the embodiment of the present application can standardize the pitch data into a unified system by counting the second pitch in the main track and standardizing it, so as to facilitate accurate analysis. The mode category and mode are confirmed based on the second frequency of the duration of each pitch, and the pitch duration frequency information is fully mined to make the judgment process more objective and scientific. This analysis method based on quantitative data can effectively capture the time distribution characteristics of the melody pitch in folk music, so as to accurately judge the mode category.
[0015] The second aspect of the present application provides a device for judging Western major and minor keys and pentatonic modes, including: a preprocessing module, used to obtain a target audio to be judged, and preprocess the target audio to obtain audio information of the target audio; an identification module, used to determine the tonic of the target audio according to the audio information, and confirm the mode category label of the target audio according to the tonic; a judgment module, used to confirm the mode of the target audio according to the mode category label, and when the mode is a Western major and minor mode or a national pentatonic mode, adopt a corresponding strategy to obtain a mode judgment result.
[0016] Through the above technical solution, the embodiment of the present application can obtain the target audio and perform preprocessing to obtain audio information, remove noise and other interference, accurately extract effective music features, and provide a high-quality data basis for subsequent accurate judgment. Comprehensively consider multiple music elements to determine the tonic based on audio information, so that the tonic judgment is more accurate and reliable; then confirm the mode category label based on the tonic, effectively narrow the mode range, and reduce the judgment workload; finally, adopt corresponding strategies for Western major and minor modes and national pentatonic modes to obtain judgment results, fully consider the uniqueness of different mode systems, and improve the accuracy and adaptability of judgment.
[0017] Optionally, in one embodiment of the present application, the preprocessing module includes: a first extraction unit, used to extract basic beat information in the music based on the target audio; a second extraction unit, used to separate the tracks in the target audio, determine the main instrument track according to the duration proportion of each track in the target audio and the uniqueness of each track, and extract the duration information of all pitches based on the main instrument track; a deduplication unit, used to retain a single melody and remove duplicates based on the duration information to obtain a non-repetitive sequence in the main track.
[0018] Through the above technical scheme, the embodiment of the present application can provide key rhythmic basis for subsequent processing by extracting basic beat information of the music, so that the entire music processing process is more in line with the rhythmic characteristics of the music itself, and the music theory basis of the scheme is enhanced. By determining the main instrument track based on the beat information, the source of the core melody can be accurately located, avoiding interference from other tracks, and ensuring that subsequent analysis focuses on key music elements. Retaining a single melody and removing duplicate operations based on duration information simplifies the data and reduces the impact of redundant information on judgment. It not only reduces the computational complexity and improves processing efficiency, but also makes the extracted non-repetitive sequence more accurately represent the essential melodic characteristics of the music, laying a solid and reliable data foundation for subsequent steps such as accurately determining the tonic and judging the mode, and improving the accuracy and effectiveness of the entire mode judgment technical scheme.
[0019] Optionally, in one embodiment of the present application, the recognition module includes: a calculation unit, used to calculate the note frequencies of the same piece of music under different rhythmic patterns according to the beat distribution of different beat types; a tuning unit, used to determine the main tone based on the note frequency statistics and its frequency after adjusting to the standardized twelve tones.
[0020] Through the above technical scheme, the embodiment of the present application can calculate the note frequency by distributing the beats according to different beat types, so that the calculation result is more in line with the actual music performance and auditory perception. The tonic is determined based on the highest frequency note and its frequency after the statistical standardization of the twelve tones, which comprehensively integrates the rhythm and pitch frequency information and avoids the limitation of judging the tonic by a single factor. When dealing with music with complex rhythm changes or irregular melody trends, the tonic can be locked more accurately, providing a solid and reliable foundation for subsequent mode judgment, and improving the accuracy and stability of the entire mode judgment process.
[0021] Optionally, in one embodiment of the present application, the judgment module includes: a first statistical unit, used to count all first pitches in the main audio track when the mode is the Western major or minor mode; a second statistical unit, used to take the tonic as a reference, standardize all first pitches to a sequence of 12 semitones above the tonic, and count the first frequency of the duration of occurrence of each pitch; a first judgment unit, used to confirm the major or minor key, key signature and tonality according to the first frequency of the duration of occurrence of each pitch, so as to obtain the mode judgment result.
[0022] Through the above technical scheme, the embodiment of the present application can count all the first pitches in the main track and standardize them, unify the pitch data in a standard framework that is easy to analyze, and make the subsequent frequency-based analysis more consistent and comparable. According to the first frequency of the duration of each pitch, key elements such as major and minor keys, key signatures and tonality are confirmed, and the pitch duration frequency information is fully utilized. This quantitative analysis method is more scientific and accurate than traditional qualitative or empirical judgment. It can capture the distribution characteristics of different pitches in the time dimension in music more meticulously, so as to accurately determine the relevant attributes of the mode, whether it is to distinguish the subtle differences between major and minor keys, or to determine the key signature and judge the tonality. It is more reliable, which greatly improves the accuracy and efficiency of Western major and minor mode judgment.
[0023] Optionally, in one embodiment of the present application, the judgment module includes: a third statistical unit, used to count all second pitches in the main track when the mode is the national pentatonic mode; a fourth statistical unit, used to take the tonic as a reference, standardize all second pitches to within a 12 semitone sequence above the tonic, and count the second frequency of the duration of occurrence of each pitch; a second judgment unit, used to confirm the mode category and mode according to the second frequency of the duration of occurrence of each pitch, so as to obtain the mode judgment result.
[0024] Through the above technical solution, the embodiment of the present application can standardize the pitch data into a unified system by counting the second pitch in the main track and standardizing it, so as to facilitate accurate analysis. The mode category and mode are confirmed based on the second frequency of the duration of each pitch, and the pitch duration frequency information is fully mined to make the judgment process more objective and scientific. This analysis method based on quantitative data can effectively capture the time distribution characteristics of the melody pitch in folk music, so as to accurately judge the mode category.
[0025] A third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for determining Western major and minor keys and pentatonic modes as described in the above embodiment.
[0026] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above method for determining Western major and minor keys and pentatonic modes.
[0027] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned method for determining Western major and minor keys and pentatonic modes.
[0028] The embodiment of the present application can obtain the target audio and pre-process it to extract audio information at the initial stage of audio processing, effectively remove noise interference, accurately refine music features, and lay a solid foundation for high-quality data for subsequent judgments. When determining the main tone, a variety of musical elements are integrated, and the note frequency is calculated based on the audio information and combined with the beat-type beat distribution. Then, the highest frequency note and its frequency after the standardized twelve tones are statistically analyzed, which overcomes the limitations of single-factor judgment, especially in complex music processing. The main tone can be accurately locked, and the accuracy and reliability of the main tone judgment are enhanced, laying a solid foundation for mode judgment. The main instrument track is determined based on the beat information to avoid interference, ensure focus on key melodic elements, and retain a single melody and de-duplication operations based on duration information, simplify data, reduce computational complexity, and improve processing efficiency. The extracted non-repetitive sequence can better reflect the essential melodic characteristics of the music. Confirming the mode category label based on the tonic can effectively narrow the scope and reduce the workload. Corresponding strategies are adopted for Western major and minor modes and national pentatonic modes. The former confirms the major and minor keys, key signatures and tonality based on frequency after statistical pitch standardization, while the latter confirms the mode category and mode based on frequency after statistical pitch standardization. This unified and highly simplified mode confirmation process fully considers the core characteristics and laws of different mode systems. Compared with the traditional complicated mode confirmation method that relies more on empirical judgment, it greatly reduces the calculation steps and data processing volume, and effectively improves the accuracy and efficiency of mode confirmation.
[0029] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0031] Figure 1 A flowchart of a method for determining major and minor keys and pentatonic keys according to an embodiment of the present application;
[0032] Figure 2 It is a structural schematic diagram of a device for determining major and minor keys and pentatonic keys according to an embodiment of the present application;
[0033] Figure 3 The figure is a structural example diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0034] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0035] The following describes the Western major and minor keys and pentatonic mode judgment method and device of the embodiment of the present application with reference to the accompanying drawings. In view of the problem that the related technologies mentioned in the above background technology are either completely focused on a certain field or generally consider all modes in the recognition between pentatonic modes and Western major and minor keys and other modes, resulting in a small range of music modes and high complexity of the recognition method, the present application provides a Western major and minor keys and pentatonic mode judgment method, in which the target audio can be obtained and pre-processed to obtain audio information, remove noise and other interference, accurately extract effective music features, and provide a high-quality data basis for subsequent accurate judgment. Comprehensively consider a variety of musical elements to determine the tonic based on audio information, so that the tonic judgment is more accurate and reliable; then confirm the mode category label based on the tonic, efficiently narrow the mode range, and reduce the judgment workload; finally, adopt corresponding strategies for Western major and minor modes and national pentatonic modes to obtain judgment results, fully consider the uniqueness of different mode systems, and improve the accuracy and adaptability of judgment. This solves the problem that the related technologies either focus entirely on a certain field or consider all modes in general when recognizing between pentatonic modes and Western major and minor modes and other modes, resulting in a small range of recognized music modes and high complexity of the recognition method.
[0036] Specifically, Figure 1 A flowchart of a method for determining major or minor key and pentatonic mode provided in an embodiment of the present application.
[0037] like Figure 1 As shown, the method for determining the major and minor keys and the pentatonic mode comprises the following steps:
[0038] In step S101, a target audio to be determined is obtained, and the target audio is preprocessed to obtain audio information of the target audio.
[0039] Optionally, in one embodiment of the present application, the target audio is preprocessed to obtain audio information of the target audio, including: extracting basic beat information in the music based on the target audio; separating the tracks in the target audio, determining the main instrument track based on the duration proportion of each track in the target audio and the uniqueness of each track, and extracting the duration information of all pitches based on the main instrument track; retaining a single melody and removing duplications based on the duration information to obtain a non-repetitive sequence in the main track.
[0040] Specifically, the basic beat information in the music is extracted. For a sung song, the vocal or main instrument track is separated from the audio file and the MIDI file is obtained, from which the duration information of all pitches is extracted. The main instrument track is determined as follows:
[0041] 1) Calculate the proportion p of the total duration of the audio track to the total audio.
[0042] 2) Use the DTW (Dynamic Time Warping) algorithm to obtain the distance matrix A between different audio tracks.
[0043] 3) Use the total duration proportion p and the row sum of the kth row after normalization of the distance matrix A The weighted average of the .
[0044] The ratio of the total duration of the tracks to the total audio is calculated and the DTW algorithm is used to obtain the distance matrix between different tracks. The weighted average of the two is then used to confirm the main instrument track. This method comprehensively considers the track duration factor and the differences in melody contours between tracks. It is fundamentally different from the traditional method of determining the main track based only on a single duration or simple pitch comparison. It can more accurately locate the core melody track in multi-track music and effectively reduce the interference data for subsequent mode analysis.
[0045] It should be noted that in music, the main track may contain some repeated melody fragments, which may interfere with the subsequent music analysis and increase unnecessary calculations. In order to simplify the data and focus on the core melody features of the music, ACF (Autocorrelation Function) is used to remove the repeated parts in the main track to obtain a non-repeating sequence. The obtained non-repeating sequence can more accurately represent the essential melody features of the music, reducing the impact of redundant information on subsequent steps such as judging the tonic and analyzing the mode.
[0046] By analyzing the main track through ACF, the non-repeating sequences in it are accurately obtained, simplifying the melody data structure. This method helps to focus on more representative melody fragments for subsequent tonic and mode confirmation, significantly reducing the complexity and amount of data processing, and improving the overall processing efficiency, which is different from the existing technology that does not perform targeted processing on the repeated parts of the main track.
[0047] In the audio preprocessing stage, the embodiment of the present application can provide an accurate data basis for the subsequent determination of the tonic based on the beat and re-beat distribution by extracting the basic beat information of the music, so that the determination of the tonic is more based on music theory and more accurate. For example, in music with complex rhythms, the beat information can effectively guide the calculation of the tonic frequency and avoid misjudgment of the tonic due to changing rhythms, which helps to improve the reliability of the entire mode recognition process. The main instrument track is determined by using the time ratio and the distance matrix rows and weighted averages obtained based on the DTW algorithm. Compared with the traditional method of judging based on time or simple pitch relationships alone, the core melody track can be separated from multi-track music more accurately. The use of the autocorrelation function to remove the repeated parts in the main track not only simplifies the melody data and reduces the complexity of data processing, but also enables the subsequent tonic and mode confirmation to focus on more representative melody fragments.
[0048] In step S102, the main tone of the target audio is determined according to the audio information, and the mode category label of the target audio is confirmed according to the main tone.
[0049] Optionally, in one embodiment of the present application, the tonic of the target audio is determined based on the audio information, including: calculating the note frequencies of the same piece of music under different rhythmic patterns according to the beat distribution of different beat types; and determining the tonic based on the note frequency statistics of the highest frequency note and its frequency after adjusting to the standardized twelve tones.
[0050] Specifically, according to the distribution of beats in different rhythmic types, the note frequencies of the same piece of music under different rhythmic types are calculated in turn, and the notes with the highest frequencies and their frequencies after adjustment to the standardized twelve tones are counted, and the tonic is determined by the maximum value.
[0051] The following uses the Western major and minor scale and pentatonic scale without considering the modulation as an example to confirm the scale category. The steps for confirming the scale category are as follows:
[0052] 1) Feature extraction: retain the pitch sequence after removing repetitions and the extracted main tone.
[0053] 2) Model construction: Build an MLP (Multilayer Perceptron) model consisting of an input layer, a hidden layer, and an output layer.
[0054] Among them, the input layer: receives n-dimensional feature vectors, where n is the feature dimension specifically selected during training;
[0055] Hidden layer: contains 2 hidden layers, using ReLU (Rectified Linear Unit) activation function;
[0056] Output layer: contains neurons of the same number of modulation types.
[0057] The forward propagation of the model can be expressed as: Y = (W2·ReLU(W1·x+b1)+b2).
[0058] Among them, W1 and W2 are the weight matrices of the first and second layers respectively, b1 and b2 are bias vectors, and x is the input feature vector. The cross entropy loss function is used to measure the difference between the predicted value and the mode category, and the Adam (Adaptive Moment Estimation) optimizer is used to minimize the loss function, and the accuracy is used to evaluate the performance of the model.
[0059] The model structure and parameter settings are obtained through experiments and optimization, and can effectively fit the complex relationship between music features and modes. Compared with traditional mode recognition methods based on simple rules or general neural network structures, it has significant improvements in mode recognition accuracy and processing efficiency, and can adapt to the mode recognition needs of different styles of music.
[0060] 3) Mode category recognition: The preprocessed audio is used as input to obtain the mode category label.
[0061] The embodiment of the present application can be an innovative method for determining the tonic in the tonic confirmation stage by calculating the note frequency based on the distribution of beats of different beat types and combining it with the standardized twelve-tone system, which is different from the previous method of relying on note duration statistics or simple inference of specific interval relationships. This idea of comprehensive consideration of rhythm and pitch frequency has obvious advantages when dealing with music with special rhythms or irregular melodies. In the mode category confirmation part, an MLP model with a specific structure is constructed for mode recognition. The combination of 2 hidden layers and the ReLU activation function it uses can relatively easily fit the complex relationship between music features and modes.
[0062] In step S103, the mode of the target audio is confirmed according to the mode category label, and when the mode is a Western major or minor mode or a national pentatonic mode, a corresponding strategy is adopted to obtain a mode determination result.
[0063] The following is an example of a Western major or minor scale or pentatonic scale without considering the modulation situation.
[0064] Optionally, in one embodiment of the present application, when the mode is a Western major or minor mode or a national pentatonic mode, a corresponding strategy is adopted to obtain a mode judgment result, including: when the mode is a Western major or minor mode, counting all the first pitches in the main audio track; taking the tonic as a reference, standardizing all the first pitches to a sequence of 12 semitones above the tonic, and counting the first frequency of the duration of occurrence of each pitch; confirming the major or minor key, key signature and tonality based on the first frequency of the duration of occurrence of each pitch to obtain a mode judgment result.
[0065] In the actual implementation process, if the mode is confirmed to be a Western major or minor mode, first count all the pitches in the main track, take the determined tonic as the benchmark, standardize all the pitches to a sequence of 12 semitones above the tonic, and count the frequency of the duration of each pitch. Secondly, calculate the interval from the first to third tone, the major third is the major key, and the minor third is the minor key. Thirdly, calculate the pitch corresponding to the tonic and get the corresponding key signature. Finally, calculate the interval between each pitch, if there is a minor third, it is a harmonic key; calculate the number of main pitches, when the number of pitches is 10, it is a melodic key; if the above two situations do not exist, it is a natural key.
[0066] Optionally, in one embodiment of the present application, when the mode is a Western major or minor mode or a national pentatonic mode, a corresponding strategy is adopted to obtain a mode judgment result, including: when the mode is a national pentatonic mode, counting all second pitches in the main track; taking the tonic as a reference, standardizing all second pitches to a sequence of 12 semitones above the tonic, and counting the second frequency of the duration of occurrence of each pitch; confirming the mode category and mode according to the second frequency of the duration of occurrence of each pitch to obtain a mode judgment result.
[0067] In the actual implementation process, if the mode is confirmed to be a national pentatonic mode, all the pitches in the main track are first counted, and the determined tonic is used as the benchmark. All pitches are standardized within a sequence of 12 semitones above the tonic, and the frequency of the duration of each pitch is counted. Secondly, the intervals between each pitch are calculated. If there are 2 minor seconds, it is a seven-tone mode; if there is 1 minor second, it is a six-tone mode; if there is no minor second, it is a pentatonic mode. Finally, the corresponding mode name is obtained based on the interval difference sequence between each note.
[0068] The embodiments of the present application can adopt a very simplified processing method for both Western major and minor modes and national pentatonic modes in the mode confirmation process. This unified and highly simplified mode confirmation process fully considers the core characteristics and laws of different mode systems. Compared with the traditional complex mode confirmation method that relies more on empirical judgment, it greatly reduces the calculation steps and data processing volume, effectively improves the accuracy and efficiency of mode confirmation, and can quickly and accurately complete the mode analysis task regardless of whether it is processing Western music or national music. It shows excellent performance advantages in scenarios such as large-scale music data processing and real-time music mode analysis, while reducing the demand and dependence on computing resources.
[0069] According to the method for judging Western major and minor keys and pentatonic modes proposed in the embodiment of this application, the target audio can be obtained and preprocessed to obtain audio information, remove noise and other interference, accurately extract effective music features, and provide a high-quality data basis for subsequent accurate judgment. Comprehensively consider various musical elements to determine the tonic based on audio information, so that the tonic judgment is more accurate and reliable; then confirm the mode category label based on the tonic, effectively narrow the mode range, and reduce the judgment workload; finally, adopt corresponding strategies for Western major and minor modes and national pentatonic modes to obtain judgment results, fully consider the uniqueness of different mode systems, and improve the accuracy and adaptability of judgment.
[0070] Next, a device for determining major or minor key and pentatonic mode according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0071] Figure 2 It is a block diagram of a device for determining major and minor keys and pentatonic modes according to an embodiment of the present application.
[0072] like Figure 2 As shown, the Western major and minor keys and pentatonic mode determination device 10 comprises: a preprocessing module 100 , a recognition module 200 and a determination module 300 .
[0073] Specifically, the preprocessing module 100 is used to obtain the target audio to be determined, and preprocess the target audio to obtain audio information of the target audio.
[0074] The identification module 200 is used to determine the main tone of the target audio according to the audio information, and confirm the mode category label of the target audio according to the main tone.
[0075] The judgment module 300 is used to confirm the mode of the target audio according to the mode category label, and adopt a corresponding strategy to obtain a mode judgment result when the mode is a Western major or minor mode or a national pentatonic mode.
[0076] Optionally, in one embodiment of the present application, the preprocessing module 100 includes: a first extraction unit, a second extraction unit and a deduplication unit.
[0077] The first extraction unit is used to extract basic beat information in the music based on the target audio.
[0078] The second extraction unit is used to determine the main instrument track according to the basic beat information, and extract the duration information of all pitches according to the main instrument track.
[0079] The de-duplication unit is used to retain a single melody and remove duplicate operations based on duration information to obtain a non-repetitive sequence in the main audio track.
[0080] Optionally, in one embodiment of the present application, the recognition module 200 includes: a calculation unit and a sound determination unit.
[0081] Among them, the calculation unit is used to calculate the note frequencies of the same music under different rhythm types according to the beat distribution of different beat types.
[0082] The tuning unit is used to determine the main tone based on the note frequency statistics and the note with the highest frequency after being adjusted to the standardized twelve tones and its frequency.
[0083] Optionally, in an embodiment of the present application, the judgment module 300 includes: a first statistical unit, a second statistical unit and a first judgment unit.
[0084] The first counting unit is used for counting all first pitches in the main audio track when the mode is a Western major or minor mode.
[0085] The second statistical unit is used for taking the tonic as a reference, standardizing all first pitches to be within a sequence of 12 semitones above the tonic, and counting the first frequency of the duration of occurrence of each pitch.
[0086] The first judgment unit is used to confirm the major and minor keys, key signatures and tonality according to the first frequencies of the durations of the appearance of each pitch, so as to obtain a mode judgment result.
[0087] Optionally, in an embodiment of the present application, the judgment module 300 includes: a third statistical unit, a fourth statistical unit and a second judgment unit.
[0088] The third statistical unit is used for counting all the second pitches in the main track when the mode is the national pentatonic mode.
[0089] The fourth statistical unit is used for taking the tonic as a reference, standardizing all the second pitches to be within a sequence of 12 semitones above the tonic, and counting the second frequency of the duration of occurrence of each pitch.
[0090] The second judgment unit is used to confirm the mode category and mode according to the second frequency of the duration of each pitch appearance to obtain a mode judgment result.
[0091] It should be noted that the above explanations of the embodiment of the method for determining the major and minor keys and pentatonic modes are also applicable to the device for determining the major and minor keys and pentatonic modes in this embodiment, and will not be repeated here.
[0092] According to the Western major and minor scale and pentatonic mode judgment device proposed in the embodiment of the present application, the target audio can be obtained and pre-processed to obtain audio information, remove noise and other interference, accurately extract effective music features, and provide a high-quality data basis for subsequent accurate judgment. Comprehensively consider a variety of musical elements to determine the tonic based on audio information, so that the tonic judgment is more accurate and reliable; then confirm the mode category label based on the tonic, effectively narrow the mode range, and reduce the judgment workload; finally, adopt corresponding strategies for Western major and minor scales and national pentatonic scales to obtain judgment results, fully consider the uniqueness of different mode systems, and improve the accuracy and adaptability of judgment.
[0093] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0094] A memory 301 , a processor 302 , and a computer program stored in the memory 301 and executable on the processor 302 .
[0095] When the processor 302 executes the program, the method for determining the major and minor keys and the pentatonic mode provided in the above embodiment is implemented.
[0096] Furthermore, the electronic device further comprises:
[0097] The communication interface 303 is used for communication between the memory 301 and the processor 302 .
[0098] The memory 301 is used to store computer programs that can be run on the processor 302 .
[0099] The memory 301 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0100] If the memory 301, the processor 302 and the communication interface 303 are implemented independently, the communication interface 303, the memory 301 and the processor 302 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0101] Optionally, in a specific implementation, if the memory 301, the processor 302 and the communication interface 303 are integrated on a chip, the memory 301, the processor 302 and the communication interface 303 can communicate with each other through an internal interface.
[0102] The processor 302 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0103] The embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above method for determining the major and minor keys and pentatonic modes.
[0104] The embodiment of the present application also provides a computer program product, including a computer program, which, when executed, is used to implement the above method for determining Western major and minor keys and pentatonic modes.
[0105] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0106] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0107] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0108] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.
[0109] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one or a combination of multiple of the following technologies known in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0110] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0111] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0112] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A method for determining major and minor keys and pentatonic keys, characterized in that: The following steps are involved: Acquire target audio to be determined, and preprocess the target audio to obtain audio information of the target audio; Determine the main tone of the target audio according to the audio information, and confirm the mode category label of the target audio according to the main tone; The mode of the target audio is confirmed according to the mode category label, and when the mode is a Western major or minor mode or a national pentatonic mode, a corresponding strategy is adopted to obtain a mode determination result.
2. The method according to claim 1, characterized in that The preprocessing of the target audio to obtain audio information of the target audio includes: Based on the target audio, extract basic beat information in the music; Separating the tracks in the target audio, determining the main instrument track according to the duration ratio of each track in the target audio and the uniqueness of each track, and extracting the duration information of all pitches according to the main instrument track; Based on the duration information, single melody is retained and repeated operations are removed to obtain a non-repeating sequence in the main audio track.
3. The method according to claim 1, characterized in that The determining the main sound of the target audio according to the audio information includes: According to the distribution of beats of different beat types, the note frequencies of the same piece of music under different rhythm types are calculated in sequence; The main tone is determined based on the note frequency statistics of the note with the highest frequency after adjustment to the standardized twelve tones and its frequency.
4. The method according to claim 1, characterized in that: When the mode is a Western major or minor mode or a national pentatonic mode, a corresponding strategy is adopted to obtain a mode determination result, including: When the mode is the Western major or minor mode, counting all first pitches in the main audio track; Taking the tonic as a reference, standardizing all first pitches to a sequence of 12 semitones above the tonic, and counting the first frequency of the duration of each pitch; The major and minor keys, key signatures and tonality are determined according to the first frequencies of the durations of occurrence of the various pitches, so as to obtain the mode determination result.
5. The method according to claim 1, characterized in that: When the mode is a Western major or minor mode or a national pentatonic mode, a corresponding strategy is adopted to obtain a mode determination result, including: When the mode is the national pentatonic mode, counting all second pitches in the main track; Taking the tonic as a reference, standardizing all the second pitches to a sequence of 12 semitones above the tonic, and counting the second frequency of the duration of each pitch; The mode category and mode are confirmed according to the second frequency of the duration of occurrence of each pitch to obtain the mode determination result.
6. A device for judging major and minor keys and pentatonic keys, characterized in that: include: A preprocessing module, used to obtain the target audio to be judged, and preprocess the target audio to obtain audio information of the target audio; An identification module, configured to determine the main tone of the target audio according to the audio information, and confirm a mode category label of the target audio according to the main tone; The judgment module is used to confirm the mode of the target audio according to the mode category label, and when the mode is a Western major or minor mode or a national pentatonic mode, adopt a corresponding strategy to obtain a mode judgment result.
7. The device according to claim 6, characterized in that The preprocessing module comprises: A first extraction unit, configured to extract basic beat information in the music based on the target audio; A second extraction unit is used to separate the tracks in the target audio, determine the main instrument track according to the duration ratio of each track in the target audio and the uniqueness of each track, and extract the duration information of all pitches according to the main instrument track; The de-duplication unit is used to retain a single melody and remove duplicates based on the duration information to obtain a non-repetitive sequence in the main audio track.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for determining the major and minor keys and pentatonic modes as described in any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method for determining the major and minor keys and pentatonic modes as described in any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that The computer program is executed to implement the method for determining the major and minor keys and pentatonic modes according to any one of claims 1 to 5.
Citation Information
Patent Citations
Chinese national pentatonic emotion recognition method and system
CN109273025A
Chinese ethnic music style identification method based on template matching
CN111081209A
Music tonality detection method and device, terminal equipment and computer storage medium
CN113012666A
Five-tone music mode identification method and system
CN114783394A
Audio processing method, computer device and computer program product
CN115862587A