Harmony Recognition and Its Model Training Method, Program Product, Device, and Storage Medium
By segmenting and feature extraction of audio data based on beat data, the problem of inaccurate positioning of chords in the existing harmony recognition methods is solved, and higher harmony recognition accuracy and lower calculation amount are achieved.
Patent Information
- Application Number
- CN202411442252.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-10-16
AI Technical Summary
The existing harmony recognition methods cannot accurately locate the chord change position because they use a fixed-length window to segment the audio signal, resulting in the harmonic recognition results being inaccurate enough.
By segmenting the training audio data based on the beat data, multiple audio segmentation segmentation segments are obtained, and feature extraction and model training are performed on these segments, and the internal parameters of the harmony recognition model are optimized to improve the accuracy of harmony recognition.
This method can more accurately characterize the chord change position of the audio data, improve the harmony recognition accuracy of the harmony recognition model, and reduce the amount of calculation in the harmony recognition process.
Smart Images

Figure CN118969012B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio processing. Specifically, it relates to a harmony recognition method, its model training method, program product, device, and storage medium. Background Art
[0002] Harmony is the multi - voice cooperation in music, which can enrich the layering and depth of music. Harmony recognition technology is used to identify different chords in music. Most of the existing harmony recognition methods first divide continuous audio signals into short - time frames through window partitioning, and then further analyze and process each frame.
[0003] Usually, a window with a fixed length is used. The signal is segmented by sliding the window on the audio signal, and harmony recognition is performed on the segmented multi - segment audio signals. Using a window with a fixed length cannot accurately locate the chord change position, and the sliding window may cross the chord transition point, easily resulting in fuzzy harmony recognition, and further leading to inaccurate harmony recognition results. Summary of the Invention
[0004] In view of this, the purpose of the embodiments of this application is to provide a harmony recognition method, its model training method, program product, device, and storage medium, so as to solve the technical problem that the harmony recognition results obtained based on the existing harmony recognition methods are not accurate enough.
[0005] In a first aspect, the embodiments of this application provide a method for training a harmony recognition model. The method includes:
[0006] Segmenting the audio data to be trained based on the beat data of the audio data to be trained to obtain a plurality of audio segmentation segments;
[0007] Extracting features from the audio segmentation segments to obtain audio feature data to be trained;
[0008] Inputting the audio feature data to be trained into the harmony recognition model to be trained, and obtaining the model recognition result output by the harmony recognition model to be trained;
[0009] Optimizing the internal parameters of the harmony recognition model to be trained according to the model recognition result and the harmony annotation result of the audio data to be trained, and obtaining a trained harmony recognition model.
[0010] In the above implementation process, the harmony recognition model training method segments the audio data to be trained based on the beat data of the audio data to be trained, obtaining multiple audio segmentation segments; extracts features from the audio segmentation segments to obtain the audio feature data to be trained; inputs the audio feature data to be trained into the harmony recognition model to be trained, obtaining the model recognition result output by the harmony recognition model to be trained; optimizes the internal parameters of the harmony recognition model to be trained according to the model recognition result and the harmony annotation result of the audio data to be trained, obtaining the trained harmony recognition model. Since this harmony recognition model segments the audio data to be trained through the beat data of the audio data to be trained, obtaining multiple audio segmentation segments, compared with the existing sliding window segmentation method, the beat data can better represent the chord change positions of the audio data to be trained, and the audio data to be trained can be segmented and located more accurately based on the beat data, thereby improving the harmony recognition accuracy of the trained harmony recognition model. The harmony recognition model obtained based on this harmony recognition model training method can obtain a higher-accuracy harmony recognition result when performing harmony recognition on the audio data to be recognized. It solves the technical problem that the harmony recognition result obtained based on the existing harmony recognition method is not accurate enough.
[0011] In addition, in the method of segmenting the audio signal using a fixed-length window, to ensure the continuity of the audio segments obtained by multiple segmentations, there is a situation of window overlap, which will bring a large amount of redundant calculations. However, this application uses a method of segmenting and locating the audio data to be trained based on the beat data, segmenting the audio data to be trained into multiple continuous and independent audio segmentation segments, which can improve the accuracy of audio segmentation and location while reducing the computational amount in the harmony recognition process.
[0012] Optionally, in the embodiment of this application, the beat data of the audio data to be trained includes the positions of stressed beats; before segmenting the audio data to be trained based on the beat data of the audio data to be trained to obtain multiple audio segmentation segments, the method further includes: performing rhythm feature recognition on the audio data to be trained to obtain the rhythm feature data of the audio data to be trained; determining the positions of the stressed beats based on the rhythm feature data.
[0013] Optionally, in the embodiment of this application, determining the positions of the stressed beats based on the rhythm feature data includes: using a period estimation method to determine the beat period of the audio data to be trained; determining the positions of the stressed beats based on the rhythm feature data and the beat period.
[0014] In the above implementation process, through the beat period and the rhythm feature data, the positions of the stressed beats of the audio data to be trained can be determined more comprehensively and accurately.
[0015] In addition, when performing audio signal segmentation using a fixed-length window, it is necessary to adjust the window length based on music with different rhythm characteristics, which increases the difficulty of harmony recognition and the labor cost. However, in this application, through the methods of rhythm feature recognition and period estimation, based on the obtained rhythm feature data and beat period, the position of the stressed beat can be accurately determined; and then, based on the position of the stressed beat, the segmentation of the audio data to be trained is realized. In the process of realizing the segmentation of the audio data to be trained provided by this application, no manual participation is required, which can reduce the labor cost while improving the accuracy of the obtained audio segmentation segments.
[0016] Optionally, in an embodiment of this application, after determining the position of the stressed beat based on the rhythm feature data, the method further includes: calculating the average beat frame of the audio data to be trained according to the position of the stressed beat; supplementing the position of the stressed beat according to the average beat frame to obtain the supplemented beat position; the segmenting the audio data to be trained based on the beat data of the audio data to be trained to obtain multiple audio segmentation segments includes: segmenting the audio data to be trained based on the supplemented beat position to obtain multiple audio segmentation segments.
[0017] In the above implementation process, by supplementing the position of the stressed beat with the average beat frame, a more complete supplemented beat position can be obtained, thereby improving the harmony recognition accuracy and precision of the harmony recognition model trained based on the audio segmentation segments.
[0018] Optionally, in an embodiment of this application, the segmenting the audio data to be trained based on the beat data of the audio data to be trained to obtain multiple audio segmentation segments includes: determining the number of beats of the audio data to be trained according to the position of the stressed beat; determining the beat speed of the audio data to be trained based on the audio length and the number of beats of the audio data to be trained; performing speed normalization processing on the audio data to be trained according to the beat speed and the standard beat speed to obtain normalized audio data; segmenting the normalized audio data based on the beat data to obtain multiple audio segmentation segments.
[0019] In the above implementation process, by performing speed normalization processing on the audio data to be trained, the sizes of multiple audio feature data to be trained obtained by feature extraction of the audio segmentation segments can be made consistent, so as to improve the harmony recognition accuracy and precision of the harmony recognition model trained based on the audio feature data to be trained.
[0020] Optionally, in the embodiments of the present application, the step of segmenting the normalized audio data based on the beat data to obtain a plurality of audio segmentation segments includes: segmenting the normalized audio data based on the beat data to obtain a plurality of initial segmentation segments; performing padding processing or truncation processing on the initial segmentation segments based on the average beat frame to obtain the processed audio segmentation segments.
[0021] In the above implementation process, the lengths of the plurality of initial segmentation segments obtained by segmenting the normalized audio data based on the beat data are not necessarily exactly the same. By performing padding processing or truncation processing on the initial segmentation segments, the slight length differences between different initial segmentation segments can be eliminated, and a plurality of audio segmentation segments with consistent lengths can be obtained, so as to improve the harmony recognition accuracy and accuracy of the harmony recognition model trained based on the audio segmentation segments.
[0022] In a second aspect, an embodiment of the present application provides a harmony recognition method, the method including:
[0023] Segmenting the audio data to be recognized based on the beat data of the audio data to be recognized to obtain a plurality of audio recognition segments;
[0024] Extracting features from the audio recognition segments to obtain audio feature data to be recognized;
[0025] Inputting the audio feature data to be recognized into a trained harmony recognition model to obtain a harmony recognition result output by the trained harmony recognition model; wherein, the harmony recognition model is trained based on the harmony recognition model training method described in any item of the first aspect above.
[0026] In the above implementation process, the harmony recognition model adopted by this harmony recognition method segments the audio data to be trained based on the beat data of the audio data to be trained to obtain a plurality of audio segmentation segments. Compared with the existing sliding window segmentation method, the beat data can better represent the chord change positions of the audio data to be trained. Based on the beat data, the segmentation and positioning of the audio data to be trained can be more accurately realized, thereby improving the harmony recognition accuracy of the trained harmony recognition model. Performing harmony recognition on the audio data to be recognized based on this harmony recognition method can obtain a harmony recognition result with higher accuracy. This solves the technical problem that the harmony recognition result obtained based on the existing harmony recognition method is not accurate enough.
[0027] In a third aspect, an embodiment of the present application further provides a computer program product, including computer programs / instructions, which when executed by a processor implement the method described in any item of the first aspect or the second aspect.
[0028] Fourthly, an embodiment of the present application further provides an electronic device; the electronic device includes:
[0029] a memory;
[0030] a processor;
[0031] A computer program executable by the processor is stored on the memory. When the computer program is executed by the processor, the method according to any one of the first aspect or the second aspect is executed.
[0032] Fifthly, an embodiment of the present application further provides a computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium. When the computer program instructions are run by a processor, the method according to any one of the first aspect or the second aspect is executed.
[0033] The beneficial effects of the present application at least include: The harmony recognition model training method divides the audio data to be trained based on the beat data of the audio data to be trained to obtain multiple audio segmentation segments; extracts features from the audio segmentation segments to obtain the audio feature data to be trained; inputs the audio feature data to be trained into the harmony recognition model to be trained to obtain the model recognition result output by the harmony recognition model to be trained; optimizes the internal parameters of the harmony recognition model to be trained according to the model recognition result and the harmony annotation result of the audio data to be trained to obtain a trained harmony recognition model. Since the harmony recognition model divides the audio data to be trained based on the beat data of the audio data to be trained to obtain multiple audio segmentation segments, compared with the existing sliding window segmentation method, the beat data can better represent the chord change position of the audio data to be trained, and the audio data to be trained can be more accurately segmented and located based on the beat data, thereby improving the harmony recognition accuracy of the trained harmony recognition model. Based on the harmony recognition model obtained by the harmony recognition model training method, performing harmony recognition on the audio data to be recognized can obtain a more accurate harmony recognition result. It solves the technical problem that the harmony recognition result obtained based on the existing harmony recognition method is not accurate enough.
[0034] Furthermore, in the method of segmenting an audio signal using a fixed-length window, in order to ensure the continuity of the audio segment signals obtained by multiple segmentations, there is a situation of window overlap, which will bring a large amount of redundant calculations. However, the present application uses a method of segmenting and positioning the audio data to be trained based on beat data, and divides the audio data to be trained into multiple continuous and independent audio segmentation segments, which can improve the accuracy of audio segmentation and positioning while reducing the calculation amount in the harmony recognition process. Description of the Drawings
[0035] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0036] Figure 1 A schematic flowchart of a method for training a harmony recognition model provided by an embodiment of the present application;
[0037] Figure 2 A schematic diagram of the position of the stressed beats of the audio data to be trained provided by an embodiment of the present application;
[0038] Figure 3 A schematic flowchart of a method for segmenting audio data provided by an embodiment of the present application;
[0039] Figure 4 A schematic flowchart of a harmony recognition method provided by an embodiment of the present application;
[0040] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0041] The following will describe in detail the embodiments of the technical solutions of the present application with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application and are only examples, and thus cannot be used to limit the protection scope of the present application.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0043] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of the present application, "a plurality of" means two or more, unless otherwise specifically defined.
[0044] Please refer to Figure 1 The schematic flowchart of a method for training a harmony recognition model provided by an embodiment of the present application shown. The method for training a harmony recognition model may include the following steps:
[0045] S101. Segment the audio data to be trained based on the beat data of the audio data to be trained to obtain a plurality of audio segmentation segments;
[0046] S102. Extract features from the audio segmentation segments to obtain the audio feature data to be trained;
[0047] S103. Input the audio feature data to be trained into the harmony recognition model to be trained, and obtain the model recognition result output by the harmony recognition model to be trained;
[0048] S104. Optimize the internal parameters of the harmony recognition model to be trained according to the model recognition result and the harmony annotation result of the audio data to be trained, and obtain the trained harmony recognition model.
[0049] Among them, in step S101, the beat data may include the beat period or the position of the stressed beat, or may include both the beat period and the position of the stressed beat. The position of the stressed beat can be represented by the time series of the occurrence of the stress in different beats. According to the beat period or the position of the stressed beat, the audio data to be trained can be segmented into multiple audio segmentation segments. The number of audio segmentation segments can be 20 segments, or 16 segments or other reasonable values; the specific number of audio segmentation segments is related to the actual application data such as the audio length and the beat period of the audio data to be trained.
[0050] Among them, in step S102, feature extraction methods such as fast Fourier transform (FFT), Mel Frequency Cepstrum Coefficient (MFCC), Constant Q Transformation (CQT), or Mel transformation can be used to extract features from the audio segmentation segments. Among them, the Mel transformation non-linearly maps the frequency axis to the Mel scale by simulating the sound perception characteristics of the human ear, so as to better realize the feature processing of the audio signal. It is also possible to select multiple feature extraction methods among FFT, MFCC, CQT, and Mel transformation to jointly extract features from the audio segmentation segments. This application does not make specific limitations on this.
[0051] Among them, in step S103, after obtaining the audio feature data to be trained, standardization processing methods such as zero-mean normalization or amplitude normalization can be used to perform standardization processing on the audio feature data to be trained, and then the standardized feature data is input into the harmony recognition model to be trained for model training; to improve the stability of model training and the harmony recognition performance.
[0052] Among them, in step S104, the corresponding loss function value can be calculated according to the model recognition result and the harmony annotation result of the audio data to be trained, and the internal parameters of the harmony recognition model to be trained are optimized based on the loss function value. Exemplarily, the loss function value between the model recognition result and the harmony annotation result can be calculated based on distance calculation methods such as Euclidean distance, cosine distance, or Manhattan distance. The harmony annotation result can be obtained by manually performing harmony annotation on the audio data to be trained.
[0053] It can be seen that the harmony recognition model training method provided by the embodiments of the present application divides the audio data to be trained based on the beat data of the audio data to be trained to obtain multiple audio segmentation segments; extracts features from the audio segmentation segments to obtain the audio feature data to be trained; inputs the audio feature data to be trained into the harmony recognition model to be trained to obtain the model recognition result output by the harmony recognition model to be trained; optimizes the internal parameters of the harmony recognition model to be trained according to the model recognition result and the harmony annotation result of the audio data to be trained to obtain a trained harmony recognition model. Since this harmony recognition model divides the audio data to be trained based on the beat data of the audio data to be trained to obtain multiple audio segmentation segments, compared with the existing sliding window segmentation method, the beat data can better represent the chord change position of the audio data to be trained, and the segmentation and positioning of the audio data to be trained can be more accurately achieved based on the beat data, thereby improving the harmony recognition accuracy of the trained harmony recognition model. Based on the harmony recognition model obtained by this harmony recognition model training method, when performing harmony recognition on the audio data to be recognized, a more accurate harmony recognition result can be obtained. It solves the technical problem that the harmony recognition result obtained based on the existing harmony recognition method is not accurate enough.
[0054] In some optional embodiments, the beat data of the audio data to be trained includes the positions of the stressed beats; before S101, dividing the audio data to be trained based on the beat data of the audio data to be trained to obtain multiple audio segmentation segments, the method further includes: performing rhythm feature recognition on the audio data to be trained to obtain the rhythm feature data of the audio data to be trained; determining the positions of the stressed beats based on the rhythm feature data.
[0055] Among them, the rhythm feature recognition of the audio data to be trained can be performed based on methods such as short-time Fourier transform or energy envelope, and the rhythm feature data of the audio data to be trained can be obtained. Among them, the energy envelope can extract the intensity change of the audio signal by analyzing the energy distribution of the audio data to be trained, and obtain the rhythm feature data of the audio data to be trained based on the intensity change of the audio signal. The short-time Fourier transform can divide the audio data to be trained into multiple short time periods (frames), and perform Fourier transform on each frame to obtain the frequency characteristics of the audio data to be trained at each time point; based on the frequency characteristics of the audio data to be trained, its rhythm feature data can be obtained. It is also possible to jointly perform rhythm feature recognition on the audio data to be trained based on the short-time Fourier transform and the energy envelope. The rhythm feature data can specifically include the measures, beats, positions of accents, and cycle periods of the audio data to be trained. According to the identified rhythm feature data, the position of the stressed beat of the audio data to be trained can be determined.
[0056] In some optional embodiments, determining the position of the stressed beat based on the rhythm feature data includes: using a period estimation method to determine the beat period of the audio data to be trained; and determining the position of the stressed beat based on the rhythm feature data and the beat period.
[0057] Among them, period estimation methods such as autocorrelation or spectral domain can be used to estimate the beat period of the audio data to be trained, and then based on the rhythm feature data and the beat period, the position of the stressed beat can be determined. The position of the stressed beat can be represented by the stress time series of different beats. Please refer to Figure 2 , Figure 2 which is a schematic diagram of the position of the stressed beat of a piece of audio data to be trained provided by an embodiment of the present application. Corresponding to Figure 2 the position of the stressed beat Tb of the audio data to be trained shown can be represented by Tb = . It should be noted that Figure 2 only exemplarily shows some positions of the stressed beats of the audio data to be trained. t0 represents the start position of the audio data to be trained, and tn represents the position of the nth stressed beat. The value of tn can be determined according to the time when the nth stressed beat appears. Since the stress does not necessarily appear in each beat, through the beat period and the rhythm feature data, the position of the stressed beat of the audio data to be trained can be determined more comprehensively and accurately.
[0058] In some optional embodiments, after determining the position of the stressed beat based on the rhythm feature data, the method further includes: calculating the average beat frame of the audio data to be trained according to the position of the stressed beat; and supplementing the position of the stressed beat according to the average beat frame to obtain the supplemented beat position.
[0059] S101. Segment the to-be-trained audio data based on the beat data of the to-be-trained audio data to obtain multiple audio segmentation segments, including: segment the to-be-trained audio data based on the supplemented beat positions to obtain multiple said audio segmentation segments.
[0060] Among them, the average beat frame can be represented by the average time difference between every two stressed beat positions; correspondingly, the average beat frame can be calculated based on ; among them, represents the average beat frame, represents the i-th stressed beat position (specifically, it can be represented by the time when the i-th stressed beat appears). According to the average beat frame, stress supplementation can be performed on places where the distance between stressed beats is too large. The stressed beat distance refers to the distance between adjacent two stressed beat positions (i.e., the time difference between the marked adjacent two stressed beats). The stressed beat positions determined based on the rhythm feature data may be incomplete (there may be missed beats). By supplementing the stressed beat positions with the average beat frame, more complete supplemented beat positions can be obtained; thereby improving the harmony recognition accuracy and accuracy of the harmony recognition model trained based on the audio segmentation segments.
[0061] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an audio data segmentation method provided by an embodiment of the present application.
[0062] In some optional embodiments, S101. Segment the to-be-trained audio data based on the beat data of the to-be-trained audio data to obtain multiple audio segmentation segments, including: S1011. Determine the number of beats of the to-be-trained audio data according to the stressed beat positions; S1012. Based on the audio length of the to-be-trained audio data and the number of beats, determine the beat speed of the to-be-trained audio data; S1013. Perform speed normalization processing on the to-be-trained audio data according to the beat speed and the standard beat speed to obtain standardized audio data; S1014. Segment the standardized audio data based on the beat data to obtain multiple said audio segmentation segments.
[0063] Among them, the number of beats m of the to-be-trained audio data can be determined according to the stressed beat positions or the stressed beat positions after stress supplementation. The beat speed of the to-be-trained audio data can be represented by bpm (i.e., the number of beats per minute of the audio data), and the beat speed of the to-be-trained audio data can be calculated based on ; among them, (Unit: seconds) represents the audio length of the audio data to be trained. The standard tempo can be 100, 120, or other reasonable values. When the tempo is less than the standard tempo, it is necessary to accelerate the audio data to be trained; when the tempo is greater than the standard tempo, it is necessary to decelerate the audio data to be trained to obtain standardized audio data. By performing speed normalization on the audio data to be trained, the sizes of multiple audio feature data to be trained obtained by extracting features from the audio segmentation segments can be made consistent, so as to improve the harmony recognition accuracy and accuracy of the harmony recognition model trained based on the audio feature data to be trained.
[0064] In some alternative embodiments, S1014. Segmenting the standardized audio data based on the beat data to obtain multiple audio segmentation segments includes: segmenting the standardized audio data based on the beat data to obtain multiple initial segmentation segments; performing padding processing or truncation processing on the initial segmentation segments based on the average beat frame to obtain the processed audio segmentation segments.
[0065] Among them, by performing padding processing or truncation processing on the initial segmentation segments, the length of the processed audio segmentation segments can be made consistent with the length of the average beat frame. The lengths of the multiple initial segmentation segments obtained by segmenting the standardized audio data based on the beat data are not necessarily exactly the same. By performing padding processing or truncation processing on the initial segmentation segments, the slight length differences between different initial segmentation segments can be eliminated, and multiple audio segmentation segments with consistent lengths can be obtained, so as to improve the harmony recognition accuracy and accuracy of the harmony recognition model trained based on the audio segmentation segments.
[0066] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a harmony recognition method provided by an embodiment of the present application. The harmony recognition method includes:
[0067] S201. Segmenting the audio data to be recognized based on the beat data of the audio data to be recognized to obtain multiple audio recognition segments;
[0068] S202. Extracting features from the audio recognition segments to obtain audio feature data to be recognized;
[0069] S203. Inputting the audio feature data to be recognized into the trained harmony recognition model to obtain the harmony recognition result output by the trained harmony recognition model; wherein, the harmony recognition model is trained based on the harmony recognition model training method described in any item of the first aspect above.
[0070] Among them, in step S201, the beat data may include the beat period or the position of the stressed beat, or may include both the beat period and the position of the stressed beat. The position of the stressed beat can be represented by the time series of the occurrence of the stress in different beats. According to the beat period or the position of the stressed beat, the audio data to be recognized can be segmented into multiple audio recognition segments. The number of audio recognition segments can be 20 segments, or 16 segments or other reasonable values; the specific number of audio recognition segments is related to data such as the audio length of the audio data to be recognized and the beat period.
[0071] Among them, in step S202, feature extraction methods such as FFT, MFCC, CQT or Mel transform can be used to extract features from the audio recognition segments. It is also possible to choose to use multiple feature extraction methods among FFT, MFCC, CQT and Mel transform to jointly extract features from the audio recognition segments.
[0072] Among them, in step S203, after obtaining the audio feature data to be recognized, standardization processing methods such as zero-mean normalization or amplitude normalization can be used to perform standardization processing on the audio feature data to be recognized, and then the standardized feature data is input into the trained harmony recognition model to obtain the harmony recognition result.
[0073] It should be noted that the specific implementation manners of steps S201 - S203 can correspond to the specific implementation manners of the above - mentioned harmony recognition model training method. Exemplarily, when step S101 specifically includes the above - mentioned steps S1011 - S1014, S201, segmenting the audio data to be recognized based on the beat data of the audio data to be recognized to obtain a plurality of audio recognition segments, may specifically include: determining the number of beats of the audio data to be recognized according to the position of the stressed beats; determining the beat speed of the audio data to be recognized based on the audio length and the number of beats of the audio data to be recognized; performing speed normalization processing on the audio data to be recognized according to the beat speed and the standard beat speed to obtain normalized audio data; and segmenting the normalized audio data based on the beat data to obtain a plurality of audio recognition segments. And the feature extraction method adopted in S202 can be the same as the feature extraction method adopted in S101. Since the harmony recognition model adopted by this harmony recognition method segments the audio data to be trained through the beat data of the audio data to be trained to obtain a plurality of audio segmentation segments, compared with the existing sliding window segmentation method, the beat data can better represent the chord change positions of the audio data to be trained, and based on the beat data, the segmentation and positioning of the audio data to be trained can be more accurately achieved, thereby improving the harmony recognition accuracy of the trained harmony recognition model. Performing harmony recognition on the audio data to be recognized based on this harmony recognition method can obtain a more accurate harmony recognition result. It solves the technical problem that the harmony recognition result obtained based on the existing harmony recognition method is not accurate enough.
[0074] The embodiment of the present application further provides a computer program product, including computer programs / instructions, which when executed by a processor, implement the harmony recognition model training method according to any one of the first aspect or the harmony recognition method according to the second aspect.
[0075] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device 300 provided by an embodiment of the present application. The electronic device 300 includes: a memory 302 and a processor 301; a computer program executable by the processor 301 is stored on the memory 302, and when the computer program is executed by the processor 301, the method according to any one of the first aspect or the second aspect is executed.
[0076] Among them, the memory 302 and the processor 301 can be interconnected and communicate with each other through a communication bus 303 and / or other forms of connection mechanisms (not shown). The memory 302 stores a computer program executable by the processor 301, and when the computer program is executed by the processor 301, the method described in the above - mentioned first aspect or the second aspect is executed.
[0077] The embodiments of the present application also provide a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by the processor 301, the methods described in the first aspect or the second aspect above are executed.
[0078] Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk or optical disk.
[0079] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed device / system and method can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment or a part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0080] In addition, in each embodiment of the embodiments of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0081] The above description is only an optional implementation manner of the embodiments of the present application. However, the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the embodiments of the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the embodiments of the present application.
Claims
1. A harmony recognition model training method, characterized in that: The method comprises: Segmenting the audio data to be trained based on the beat data of the audio data to be trained to obtain a plurality of audio segmentation segments; Extracting features from the audio segmentation fragments to obtain audio feature data to be trained; Inputting the audio feature data to be trained into the harmony recognition model to be trained, and obtaining a model recognition result output by the harmony recognition model to be trained; According to the model recognition result and the harmony labeling result of the audio data to be trained, optimizing the internal parameters of the harmony recognition model to be trained to obtain a trained harmony recognition model; Wherein, the beat data of the audio data to be trained includes the stress beat position; before segmenting the audio data to be trained based on the beat data of the audio data to be trained to obtain a plurality of audio segmentation segments, the method further includes: performing rhythm feature recognition on the audio data to be trained to obtain the rhythm feature data of the audio data to be trained; and determining the stress beat position based on the rhythm feature data; After determining the stress beat position based on the rhythm feature data, the method further includes: calculating an average beat frame of the audio data to be trained according to the stress beat position; supplementing the stress beat position according to the average beat frame to obtain a supplemented beat position; The step of segmenting the audio data to be trained based on the beat data of the audio data to be trained to obtain a plurality of audio segmentation segments includes: segmenting the audio data to be trained based on the supplemented beat positions to obtain a plurality of audio segmentation segments; The method of segmenting the audio data to be trained based on the beat data of the audio data to be trained to obtain a plurality of audio segmentation segments includes: determining the number of beats of the audio data to be trained according to the stress beat position; and determining the number of beats of the audio data to be trained based on the audio length of the audio data to be trained, the number of beats and the Determine the tempo of the audio data to be trained; perform tempo normalization processing on the audio data to be trained according to the tempo and the standard tempo to obtain standardized audio data; segment the standardized audio data based on the tempo data to obtain a plurality of audio segmentation segments; m represents the audio length of the audio data to be trained, m represents the number of beats, and bpm represents the beat speed.
2. The method according to claim 1, characterized in that The step of determining the stress beat position based on the rhythm feature data comprises: Using a cycle estimation method to determine the beat cycle of the audio data to be trained; The accent beat position is determined based on the rhythm feature data and the beat period.
3. The method according to claim 1, characterized in that The step of segmenting the standardized audio data based on the beat data to obtain a plurality of audio segmentation segments includes: Segmenting the standardized audio data based on the beat data to obtain a plurality of initial segmented segments; The initial segmented segments are padded or truncated based on the average beat frame to obtain the processed audio segmented segments.
4. A harmony recognition method, characterized in that: The method comprises: Segmenting the audio data to be recognized based on the beat data of the audio data to be recognized to obtain a plurality of audio recognition segments; Extracting features from the audio recognition segment to obtain audio feature data to be recognized; The audio feature data to be identified is input into a trained harmony recognition model to obtain a harmony recognition result output by the trained harmony recognition model; wherein the harmony recognition model is trained based on the harmony recognition model training method as described in any one of claims 1 to 3 above.
5. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 4 is implemented.
6. An electronic device, characterized in that: The electronic device comprises: Memory; processor; The memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the method according to any one of claims 1 to 4 is performed.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 4 is executed.
Citation Information
Patent Citations
Chord recognition method combining SVM with enhanced PCP
CN103714806A
Music editing method and recording medium which records the method
JP2001296866A
Display control device and program
JP2014056485A