Audio music score comprehensive analysis method based on deep learning

Through the comprehensive analysis method of audio scores based on deep learning, the problem of insufficient accuracy of audio conversion into music scores in the prior art is solved, and higher matching accuracy and recognition processing optimization effects are achieved.

CN119993207AInactive Publication Date: 2025-05-13HUIZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510158476.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The lack of problems in the prior art to determine whether the identification process is qualified based on the number of notes in the music score and the notes in the audio, and to optimize the identification process based on the comparison results, resulting in insufficient accuracy of converting the audio into the music score.

Method used

The comprehensive analysis method of audio scores based on deep learning is adopted. By setting the training set, training audio information and training score information are obtained, pre-processing, segmentation processing and recognition processing are performed, and recognition score information is matched with the recognition score information and the training score information, and the instructions are generated to optimize the recognition processing process.

Benefits of technology

It improves the accuracy of the conversion of audio into music scores, ensures the matching of music scores and audio, optimizes the recognition processing process, reduces the number of abnormal notes, and improves the pass rate of recognition processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993207A_ABST
    Figure CN119993207A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio analysis, in particular to an audio music score comprehensive analysis method based on deep learning. According to the method, training audio information and training music score information in a set training set are obtained, the training audio information is preprocessed to obtain noise reduction audio information, and then the noise reduction audio information is segmented to obtain a plurality of segmented audio information; then performing independent identification processing on the plurality of segmented audio information to obtain identification music score information with a plurality of identification notes, and matching the identification music score information with the training music score information to determine the number of the identification notes which can be matched with the training notes in the training music score information in the identification music score information; and generating a corresponding instruction based on the number to determine the division segment number of the noise reduction audio information, and identifying the beat number, the pickup sensitivity or the noise reduction multiplying power of the music score information, thereby optimizing an audio music score identification system or optimizing an identification model, and improving the accuracy in the process of converting the audio into the music score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio analysis, and in particular to a method for comprehensive analysis of audio scores based on deep learning. Background Art

[0002] With the development of science and technology, especially the continuous advancement of artificial intelligence technology, it has become possible to convert audio into music scores. For music learners, music scores are an important tool for learning music. By converting audio into music scores, learners can understand the melody and rhythm of music more conveniently and deeply, thereby quickly improving their musical literacy.

[0003] Regarding the technology of converting audio into music scores, the Chinese patent publication number in the prior art: CN108986841 B discloses an audio information processing method, device and storage medium, the method comprising: obtaining audio data, analyzing and processing the audio data, determining audio parameters corresponding to the audio data, and then obtaining the music score corresponding to the audio data according to the audio parameters and the audio parameters of the preset standard sound. However, this technical solution only obtains the music score through the audio parameters and the audio parameters of the preset standard sound, and does not involve how to determine whether the recognition process is qualified, and how to optimize the recognition process to improve the conversion accuracy. Therefore, the process of converting audio into music scores lacks the ability to improve itself, and it cannot guarantee that the converted music score can accurately correspond to the audio before the conversion. Summary of the invention

[0004] To this end, the present invention provides an audio score comprehensive analysis method based on deep learning, so as to solve the problem in the prior art of lacking a method for determining whether a recognition process is qualified based on the number of matches between notes in the score and notes in the audio, and optimizing the recognition process based on the comparison results.

[0005] To achieve the above object, the present invention provides an audio score comprehensive analysis method based on deep learning, comprising:

[0006] Setting a training set, and obtaining training audio information and training music score information in a single training set;

[0007] Preprocessing the training audio information to obtain noise reduction audio information;

[0008] Performing segment processing on the noise reduction audio information to obtain a plurality of segmented audio information;

[0009] Performing separate recognition processing on the plurality of segmented audio information to obtain recognized music score information with a plurality of recognized notes;

[0010] matching the identified musical score information with the training musical score information to determine the number of training notes in the training musical score information corresponding to the identified musical notes;

[0011] Based on the number, a corresponding instruction is generated to determine the number of divided segments of the noise reduction audio information, the number of beats of the recognized music score information, the sound pickup sensitivity or the noise reduction ratio.

[0012] Further, the process of generating a corresponding instruction based on the number includes: determining whether the recognition processing of the training audio information is qualified based on a comparison result of the note matching ratio and a preset note matching ratio or based on the total duration of the training audio information, wherein the note matching ratio is a ratio between the number of recognized notes matching the training notes and the total number of training notes;

[0013] When determining failure, the cause of failure is determined based on the ratio of the number of beats and the number of tones of several abnormal notes, wherein the abnormal notes are the recognized notes that do not match the training notes.

[0014] Furthermore, the process of re-determining whether the recognition processing of the training audio information is qualified based on the total duration of the training audio information includes: determining the reason for failure based on the determination result, or correcting the number of divided segments based on the total number of training notes.

[0015] Furthermore, the number of divided segments is determined based on a comparison result of the total number of training notes with a preset total number of training notes, and the number of divided segments is directly proportional to the total number of training notes.

[0016] Further, the process of determining the reason why the recognition processing of the training audio information is unqualified includes: determining the reason for the unqualified based on the comparison result of the number of beat mismatches and the number of pitch mismatches, and generating corresponding instructions based on the determined reason to determine the number of beats, sound pickup sensitivity or noise reduction ratio of the recognized music score;

[0017] The beat mismatch number is the number of abnormal notes whose recognized beats do not match the training beats, and the pitch mismatch number is the number of abnormal notes whose recognized pitches do not match the training pitches.

[0018] Further, the process of determining the beat count of the identified music score based on the corresponding instruction includes: sequentially calculating the beat count difference between a single training note and the corresponding abnormal note, and drawing a time-beat count difference curve based on the beat count difference;

[0019] Calculating the slope variance based on the curve, and determining the number of beats of the identified music score information based on a comparison result of the slope variance with a preset slope variance, or redetermining the reason based on an average slope;

[0020] Wherein, the slope of the curve corresponding to the abnormal note is obtained based on the curve, and the average slope is the average value between the slopes of each curve.

[0021] Further, the process of re-determining the cause based on the average slope includes: determining whether to correct the beat number of the recognized music score information based on a comparison result of the average slope with a preset average slope.

[0022] Further, in the case of determining to correct the beat number of the identified music score information based on the average slope, it is determined based on a comparison result of the average slope with zero whether to increase the beat number of the identified music score information or decrease the beat number of the identified music score information, and the amplitude of the increase or decrease is proportional to the absolute value of the average slope.

[0023] Furthermore, the process of determining the sound pickup sensitivity based on the corresponding instruction includes: determining to increase the sound pickup sensitivity based on a comparison result of the abnormality ratio and a preset abnormality ratio, and the increase range of the sound pickup sensitivity is proportional to the abnormality ratio;

[0024] The abnormal proportion is the ratio between the number of abnormal notes and the total number of recognized notes.

[0025] Further, the process of determining the noise reduction ratio based on the corresponding instruction includes: determining to increase the noise reduction ratio based on the comparison result of the abnormal tone ratio and the preset abnormal tone ratio, and the increase amplitude of the noise reduction ratio is proportional to the abnormal tone ratio;

[0026] The abnormal notes with mismatched pitch are recorded as abnormal pitch notes, and the abnormal pitch ratio is the ratio between the number of abnormal pitch notes and the total number of recognized notes.

[0027] Compared with the prior art, the beneficial effect of the audio score comprehensive analysis method based on deep learning of the present invention is that the method obtains training audio information and training score information in a set training set, and pre-processes the training audio information to obtain noise reduction audio information, then segment the noise reduction audio information to obtain a number of segmented audio information, and then individually recognize the several segmented audio information to obtain recognized score information with a number of recognized notes, matches the recognized score information with the training score information to determine the number of recognized notes in the recognized score information that can match the training notes in the training score information, generates corresponding instructions based on the number to determine the number of segments of the noise reduction audio information, the number of beats of the recognized score information, the pickup sensitivity or the noise reduction ratio, and then optimizes the audio score recognition system or optimizes the recognition model, thereby improving the accuracy of the audio conversion into score.

[0028] Furthermore, the present invention determines whether the recognition processing process for the training audio information is qualified based on the comparison result between the note matching ratio and the preset note matching ratio, thereby being able to quickly determine the current recognition processing result for the training audio information.

[0029] Furthermore, the present invention also re-judges the recognition processing process of the training audio information by comparing the total duration of the training audio information with the preset total duration of the training audio information, thereby improving the accuracy of the judgment process and reducing misjudgments.

[0030] Furthermore, the present invention also accurately determines the cause of failure based on the ratio between the number of beats and the number of tones of the abnormal notes, and generates corresponding instructions based on the determined cause to determine the number of beats, pickup sensitivity or noise reduction factor of the identified music score, thereby optimizing the corresponding steps in the recognition process, thereby reducing the number of abnormal notes and improving the pass rate of the recognition process.

[0031] Furthermore, the present invention determines the beat number of the identified music score information according to the comparison result of the slope variance and the preset slope variance, and then adjusts the beats in the abnormal notes, thereby reducing the number of abnormal notes.

[0032] Furthermore, the present invention also determines to increase the sound pickup sensitivity according to the comparison result of the abnormal proportion and the preset abnormal proportion, thereby improving the conversion efficiency of the electrical signal and the sound signal, thereby reducing the number of abnormal notes.

[0033] Furthermore, the present invention also determines to increase the noise reduction ratio according to the comparison result of the abnormal pitch ratio with the preset abnormal pitch ratio, and then adjusts the pitch in the abnormal note, thereby reducing the number of abnormal notes.

[0034] Furthermore, the present invention further re-determines a corresponding processing method based on the average slope to reduce the error generated when making a determination based on the slope variance.

[0035] Furthermore, the present invention determines whether to increase or decrease the beat number of the identified music score information based on the comparison result of the average slope with zero after the beat number correction of the identified music score information is completed, thereby correcting the beat number of the identified music score information and reducing the number of abnormal notes. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Schematic diagram of the module of the audio score comprehensive analysis method based on deep learning in this embodiment;

[0037] Figure 2 Schematic diagram of the process of the audio score comprehensive analysis method based on deep learning in this embodiment;

[0038] Figure 3 is a logic flow chart for determining whether the recognition processing for the training audio information is qualified based on the note matching ratio in this embodiment;

[0039] Figure 4 A logic flow chart of a processing method for determining the reason for the unqualified recognition processing of the training audio information based on the comparison result of the beat mismatch number and the tone mismatch number in this embodiment;

[0040] Figure 5 It is a logic flow chart for determining the beat number of the identified music score information based on the slope variance or re-determining the reason based on the average slope in this embodiment. DETAILED DESCRIPTION

[0041] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0043] It should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the term "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0044] In this embodiment, a method for comprehensive analysis of audio and music scores based on deep learning is provided, which can be applied to audio analysis operations. It can determine whether the recognition processing process for training audio information is qualified based on the number of abnormal notes identified, and when it is determined to be unqualified, it can determine the corresponding cause and generate a corresponding processing method, so as to optimize the recognition processing steps or optimize the recognition model, thereby improving the qualified rate in the recognition processing process and improving the recognition processing accuracy, so that the device using the method can quickly and accurately convert audio into music scores. At the same time, improving the accuracy of converting audio into music scores is also helpful for music research and analysis.

[0045] See also Figure 1As shown, it is a system that can apply the audio score comprehensive analysis method based on deep learning in this embodiment. The system includes: a training set, a data acquisition module, a data preprocessing module, a data segmentation module, a data identification module, a data matching module, an analysis module, and an adjustment module. The data acquisition module is connected to the training set to obtain the training audio information and training score information in a single training set, wherein the training score information includes a number of training notes and training beats and training tones corresponding to the training notes; the data preprocessing module is connected to the data acquisition module to obtain the training audio information and preprocess the training audio information to obtain noise reduction audio information; the data segmentation module is connected to the data preprocessing module to obtain the noise reduction information and output the noise reduction information in segments to obtain a number of segmented audio information; the data identification module is connected to the data segmentation module to obtain a number of the segmentation modules and perform separate identification processing on each segmentation module in turn to obtain the identification score information with a number of identified notes, wherein the identification score information also includes the identification beat and identification tone corresponding to the identified notes; The data matching module is connected to the data acquisition module and the data identification module respectively, and is used to match the training score information and the identification score information to determine the number of identification notes corresponding to the training notes, wherein matching means that the training beat and the identification beat, as well as the training tone and the identification tone can correspond; the analysis module is connected to the data identification module, and is used to count the number of identification notes corresponding to the notes, determine whether the identification processing for the training audio information is qualified based on the number, and generate corresponding instructions based on the reasons for failure; the adjustment module is connected to the analysis module, the data preprocessing module, the data segmentation module and the data identification module respectively, and is used to determine the number of segments of the noise reduction audio information, the number of beats of the identification score information, the sound pickup sensitivity or the noise reduction ratio based on the instructions. The training notes in the training score information correspond to the training notes in the training audio information one by one. The method continuously optimizes the identification processing process or optimizes the identification model to improve the accuracy of the device using the method for converting audio into music scores, thereby ensuring that the subsequent device using the method can be generally applicable to the function of converting audio into music scores and can ensure the conversion accuracy.

[0046] See also Figure 2 Specifically, in order to improve or optimize the function of the audio score analysis system for recognizing notes in audio, this embodiment provides an audio score comprehensive analysis method based on deep learning, and the process of the method is as follows:

[0047] S1: Setting a training set, and obtaining training audio information and training music score information in a single training set.

[0048] S2: Preprocess the training audio information to obtain noise reduction audio information.

[0049] S3: Output the noise reduction audio information in segments to obtain a plurality of segmented audio information.

[0050] S4: performing separate recognition processing on the plurality of segmented audio information to obtain recognized music score information with a plurality of recognized musical notes.

[0051] S5: Determine the number of training notes in the training music score information corresponding to the recognized notes based on matching the recognized music score information with the training music score information.

[0052] S6: Generate corresponding instructions based on the number to determine the number of segments of the noise reduction audio information, the number of beats of the recognized music score information, the sound pickup sensitivity or the noise reduction ratio.

[0053] Specifically, multiple training sets are set, and the number of data samples is ensured to meet the test requirements by identifying audio information for multiple training sets, so that the corresponding system can obtain a large amount of data support; the training music score information includes several training notes. First, the training audio information is preprocessed to obtain the corresponding noise reduction audio information, wherein the preprocessing method includes data cleaning, and the noise and redundant information in the training audio information are removed by data cleaning to ensure the purity and accuracy of the data. Specifically, the background noise in the training audio information can be removed by a filter, thereby improving the subsequent recognition processing accuracy. Then, the obtained noise reduction audio information is segmented to obtain segmented audio information after several segments, so that the relatively long training audio information can be divided into multiple relatively short training audio information, and each relatively short training audio information is identified separately in turn, thereby reducing the subsequent recognition processing pressure, wherein, in this embodiment, the duration of each segmented audio information is the same, and in other embodiments, the duration of each segmented audio information can also be different. Then, each segmented audio information is individually identified and processed to obtain segmented identified music score information with a number of identified notes, and then each segmented identified music score information is integrated into identified music score information, wherein the identification process can also identify the beat and tone (or pitch) corresponding to the identified notes. The training music score information includes a number of training notes, the training music score information and the identified music score information are matched, and the number of identified notes in the identified music score information that can correspond to the training notes in the training music score information is calculated. The number can be used to determine whether the identification process for the training audio information is qualified. If qualified, the data collected in this identification process is transmitted to the above system, so that the system records the corresponding parameters in this identification process. If unqualified, the reason for the unqualified is determined based on the number of note matches in this identification process. Based on this reason, a corresponding processing method is generated, including determining the number of segments of the noise reduction audio information, the number of beats for identifying the music score information, the sound pickup sensitivity or the noise reduction ratio. The corresponding processing method in the method is used to optimize the system or the deep learning model step by step, and then a system or model is trained to specifically recognize and convert audio into music scores, and the accuracy of the audio conversion process is improved.

[0054] Furthermore, the process of generating corresponding instructions based on the quantity includes: determining whether the recognition processing of the training audio information is qualified based on the comparison result of the note matching ratio and the preset note matching ratio or based on the total duration of the training audio information, wherein the note matching ratio is the ratio between the number of recognized notes that match the training notes and the total number of training notes; when it is determined to be unqualified, determining the reason for the unqualified processing based on the ratio between the number of beats and the number of tones of several abnormal notes, wherein the abnormal notes are the recognized notes that do not match the training notes.

[0055] Specifically, in this embodiment, the beat of the abnormal note does not match the beat of the training note or the pitch of the abnormal note does not match the pitch of the training note. When it is determined that the recognition of the training audio information is unqualified, the reason for the unqualified can be determined based on the ratio of the number of beats and the number of pitches of several abnormal notes; the preset note matching ratio G0 can be divided into a first preset note matching ratio G1 and a second preset note matching ratio G2, and the preset note matching ratio standard G3=0.99, G1=0.96×G3, G2=G3 is set. It should be noted that G1, G2 and G3 can also be set to other values, and subsequently assigned as needed to adjust the values ​​to improve the recognition processing accuracy, thereby optimizing the recognition processing process, and the values ​​are rounded up; the comparison process based on the note matching ratio G with G1 and G2 is as follows:

[0056] If G≤G1, it means that the number of matches between the current recognized notes and the training notes is relatively small, so it can be determined that the recognition process for the training audio information is unqualified. In this embodiment, the number of abnormal notes is counted, and the number of abnormal notes with mismatched beats and the number of abnormal notes with mismatched pitches are further counted. The reason for the unqualified recognition process is determined based on the ratio between the number of mismatched beats and the number of mismatched pitches. If G1<G≤G2, it is impossible to accurately determine whether the current recognition process for the training audio information is qualified based on the number of training notes that the recognized notes can match at this time. A secondary determination is required to ensure the accuracy of the determination result. At this time, a new determination can be made based on the total duration of the training audio information. If G2<G, it means that the number of matches between the current recognized notes and the training notes is relatively large, which meets the requirements for audio conversion into music scores, so it can be determined that the recognition process for the training audio information is qualified.

[0057] Furthermore, the process of re-determining whether the recognition processing of the training audio information is qualified based on the total duration of the training audio information includes: determining the reason for failure based on the determination result, or correcting the number of divided segments based on the total number of training notes.

[0058] Specifically, in this embodiment, the reason for failure is determined according to the comparison result of the total duration of the training audio information with the preset total duration of the training audio information, or the number of segments of the noise reduction audio information is increased; the preset total duration W0 of the training audio information is set to 4 minutes, and it should be noted that W0 can also be assigned to other values, which can be set as needed; the comparison process based on the total duration W of the training audio information and W0 is as follows:

[0059] If W≤W0, it means that the total duration of the current training audio information is relatively short, so the number of segments for the noise reduction audio information based on the duration of the current training audio information will not affect the recognition process, and at this time G1<G≤G2, it can be determined that the recognition process for the training audio information is unqualified, and the reason for the unqualified needs to be determined based on the ratio between the number of beat mismatches and the number of pitch mismatches. If W0<W, it means that the total duration of the current training audio information is relatively long, so the number of segments for the noise reduction audio information based on the duration of the current training audio information will affect the recognition process, so that the number of recognized notes that do not match the training notes increases, so that G1<G≤G2 appears. At this time, it is necessary to increase the number of segments for the noise reduction audio information, thereby increasing the number of segments of the segmented audio information. At this time, the duration of each segmented audio information is relatively shortened, so that the time required for the recognition process of each segmented audio information is shortened, so as to achieve the effect of improving the accuracy of the recognition process.

[0060] Furthermore, the number of divided segments is determined based on a comparison result of the total number of training notes with a preset total number of training notes, and the number of divided segments is directly proportional to the total number of training notes.

[0061] Specifically, in this embodiment, the number of division segments is increased according to the comparison result of the total number of training notes with the preset total number of training notes; the preset total number of training notes T0 can be divided into a first preset total number of training notes T1 and a second preset total number of training notes T2, and the preset total number of training notes is set to T3=1000 beats, T1=1.2×T3, T2=1.35×T3, and it should be noted that T1, T2 and T3 can all be assigned to other values; the comparison process based on the total number of training notes T with T1 and T2 is as follows:

[0062] If T3<T≤T1, the first segment number adjustment coefficient is used to increase the number of divided segments to twice the initial value; if T1<T≤T2, the second segment number adjustment coefficient is used to increase the number of divided segments to three times the initial value; if T2<T, the third segment number adjustment coefficient is used to increase the number of divided segments to four times the initial value; in this embodiment, the increase rate of the number of divided segments can also be other values ​​and they are all integers.

[0063] Furthermore, the process of determining the reason why the recognition processing of the training audio information is unqualified includes: determining the reason for the failure based on the comparison result of the beat mismatch number and the pitch mismatch number, and generating corresponding instructions based on the determined reason to determine the beat number, pickup sensitivity or noise reduction factor of the recognized music score; wherein the beat mismatch number is the number of abnormal notes whose recognized beat does not match the training beat, and the pitch mismatch number is the number of abnormal notes whose recognized pitch does not match the training pitch.

[0064] Specifically, in this embodiment, the training score information also includes the training beat and the training tone corresponding to the training notes, and the recognition score information also includes the recognition beat and the recognition tone corresponding to the recognition notes, wherein the recognition beat and the recognition tone correspond to the recognition notes, and the training beat and the training tone correspond to the training notes; the comparison process based on the ratio of the beat mismatch number Ma to the tone mismatch number Mb is as follows:

[0065] If Ma / Mb>1.1, it means that the number of beat mismatches in the current abnormal note is greater than the number of pitch mismatches, indicating that the training audio did not meet the expected standard during recognition, resulting in a relatively large number of beat mismatches. At this time, it can be determined that the reason for the unqualified recognition process is that the recognition process of the training audio information does not meet the standard. The beats in the abnormal note can be adjusted by correcting the number of beats in the recognized music score. If 0.9≤Ma / Mb≤1.1, it means that the number of beat mismatches in the current abnormal note is not much different from the number of pitch mismatches. When the number of recognized notes and training notes differs greatly at this time, it can be determined that the reason for the unqualified recognition process is that the collection process of the training audio information does not meet the standard. The collection accuracy of the training audio information can be improved by increasing the pickup sensitivity. In this embodiment, increasing the pickup sensitivity mainly increases the sensitivity value, thereby improving the efficiency of the conversion between electrical signals and sound signals. If Ma / Mb<0.9, it means that the number of mismatched tones in the current abnormal note is greater than the number of mismatched beats, indicating that the preprocessing of the training audio information during noise reduction does not meet the expected standard, resulting in a relatively large number of mismatched tones. At this time, it can be determined that the reason for the unqualified recognition process is that the noise reduction preprocessing of the training audio information does not meet the standard. The noise reduction factor can be appropriately increased to adjust the tones in the abnormal notes. In this embodiment, the ratio range of Ma to Mb can also be other numerical ranges, which are not specifically limited here.

[0066] Furthermore, the process of determining the beat count of the identified music score based on the corresponding instructions includes: calculating the beat count difference between a single training note and the corresponding abnormal note in sequence, and drawing a time-beat count difference curve based on the beat count difference; calculating the slope variance based on the curve, and determining the beat count of the identified music score information based on a comparison result of the slope variance and a preset slope variance, or redetermining the cause based on the average slope; wherein, the slope of the curve corresponding to the abnormal note is obtained based on the curve, and the average slope is the average of the slopes of each curve.

[0067] Specifically, in this embodiment, any recognized note corresponds to one beat, and the beat difference is the number of beats that differ between the abnormal note and the corresponding training note when the beats do not match; the preset slope variance H0 is set to 0.6; the comparison process based on the slope variance H and H0 is as follows:

[0068] If H≤H0, it means that the number of beats between the abnormal notes and the training notes in the current recognition score information due to the beat mismatch is relatively small, so it can be determined that the recognition processing of the training audio information is delayed. At this time, the number of beats in each measure of the recognition score information can be corrected based on the training score information to reduce the beat difference. If H0<H, it means that the number of beats between the abnormal notes and the training notes in the current recognition score information due to the beat mismatch is relatively large. At this time, a new judgment can be made based on the average slope P to determine the corresponding processing method.

[0069] Further, the process of re-determining the cause based on the average slope includes: determining whether to correct the beat number of the recognized music score information based on a comparison result of the average slope with a preset average slope.

[0070] Specifically, in this embodiment, the preset average slope P0 is set to [-30, 30], and the comparison process based on the average slope P and P0 is as follows:

[0071] If P∈P0, it means that the absolute value of the current P is relatively small, so it can be determined that the recognition process of the training audio information is delayed. At this time, the number of beats in each measure of the recognized music score information can be corrected based on the training music score information. This means that the absolute value of the current P is relatively large, so it can be determined that an error has occurred in the construction of the recognition model for the recognition process, and an update notification for the recognition model needs to be issued.

[0072] Further, when determining to correct the beat number of the identified music score information based on the average slope, it is determined based on a comparison result of the average slope with zero whether to increase the beat number of the identified music score information or decrease the beat number of the identified music score information, and the increase or decrease amplitude is proportional to the absolute value of the average slope.

[0073] Specifically, in this embodiment, the comparison process based on the average slope P and 0 is as follows:

[0074] If P < 0, it means that the difference in the number of beats corresponding to each abnormal note from the beginning to the end of the slope curve is gradually decreasing. At this time, the number of beats corresponding to the abnormal notes on the identified music score can be reduced in turn, thereby correcting the number of beats of the identified music score information and improving the accuracy of the recognition processing for the training audio. If 0 < P, it means that the difference in the number of beats corresponding to each abnormal note from the beginning to the end of the slope curve is gradually increasing. At this time, the number of beats corresponding to the abnormal notes on the identified music score can be increased in turn, thereby correcting the number of beats of the identified music score information, reducing the number of abnormal notes and improving the accuracy of the recognition processing for the training audio. In this embodiment, the degree of increase or decrease in the number of beats is proportional to the absolute value of the average slope P. The larger the absolute value of P, the larger the difference in the number of beats, and therefore the greater the degree of increase or decrease in the number of beats in the correction.

[0075] Furthermore, the process of determining the sound pickup sensitivity based on the corresponding instruction includes: determining to increase the sound pickup sensitivity based on the comparison result of the abnormality ratio and the preset abnormality ratio, and the increase in the sound pickup sensitivity is proportional to the abnormality ratio; wherein the abnormality ratio is the ratio between the number of abnormal notes and the total number of identified notes.

[0076] Specifically, in this embodiment, the preset abnormality proportion E0 can be divided into a first preset abnormality proportion E1 and a second preset abnormality proportion E2, and the preset abnormality proportion standard E3=0.08, E1=0.5×E3, E2=1.5×E3, it should be noted that E1, E2 and E3 can all be set to other values, and all are rounded up without specific limitation; the comparison process based on the abnormality proportion E with E1 and E2 is as follows:

[0077] If E≤E1, the first sound pickup sensitivity adjustment coefficient is used to increase the sound pickup sensitivity to 1.35 times the initial value; if E1<E≤E2, the second sound pickup sensitivity adjustment coefficient is used to increase the sound pickup sensitivity to 1.55 times the initial value; if E2<E, the third sound pickup sensitivity adjustment coefficient is used to increase the sound pickup sensitivity to 1.85 times the initial value; in this embodiment, the increase rate of the sound pickup sensitivity can also be other values.

[0078] Furthermore, the process of determining the noise reduction ratio based on the corresponding instructions includes: determining to increase the noise reduction ratio based on a comparison result of the pitch abnormality ratio with a preset pitch abnormality ratio, and the increase in the noise reduction ratio is proportional to the pitch abnormality ratio; wherein the abnormal notes with mismatched pitch are recorded as pitch abnormality notes, and the pitch abnormality ratio is the ratio between the number of pitch abnormal notes and the total number of identified notes.

[0079] Specifically, in this embodiment, the preset tone abnormality ratio K0 can be divided into a first preset tone abnormality ratio K1 and a second preset tone abnormality ratio K2, and the preset tone abnormality ratio K3 is set to 0.05, K1 = 0.6 × K3, K2 = 1.3 × K3. It should be noted that K1, K2 and K3 can all be set to other values, and all are rounded up without specific limitation; the comparison process based on the tone abnormality ratio K with K1 and K2 is as follows:

[0080] If K≤K1, the first noise reduction ratio adjustment coefficient is used to increase the noise reduction ratio to twice the initial value; if K1<K≤K2, the second noise reduction ratio adjustment coefficient is used to increase the noise reduction ratio to three times the initial value; if K2<K, the third noise reduction ratio adjustment coefficient is used to increase the noise reduction ratio to four times the initial value; in this embodiment, the increase value of the noise reduction ratio can also be other values.

[0081] Furthermore, if the recognition process of the training audio information is still unqualified when the noise reduction ratio is increased, the distance between the output end of the training audio information and the receiving end of the audio recognition processing is shortened to ensure the quality of the training audio information received by the receiving end. At the same time, the audio volume of the output end is lowered to reduce noise interference.

[0082] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0083] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for comprehensive analysis of audio scores based on deep learning, characterized in that: include: Setting a training set, and obtaining training audio information and training music score information in a single training set; Preprocessing the training audio information to obtain noise reduction audio information; Performing segment processing on the noise reduction audio information to obtain a plurality of segmented audio information; Performing separate recognition processing on the plurality of segmented audio information to obtain recognized music score information with a plurality of recognized notes; matching the identified musical score information with the training musical score information to determine the number of training notes in the training musical score information corresponding to the identified musical notes; A corresponding instruction is generated based on the number to determine the number of segments of the noise reduction audio information, the number of beats of the recognized music score information, the sound pickup sensitivity or the noise reduction ratio.

2. The method for comprehensive analysis of audio scores based on deep learning according to claim 1, characterized in that: The process of generating corresponding instructions based on the quantity includes: Determining whether the recognition processing of the training audio information is qualified based on a comparison result of the note matching ratio and a preset note matching ratio or based on the total duration of the training audio information, wherein the note matching ratio is a ratio between the number of recognized notes matching the training notes and the total number of training notes; When determining failure, the cause of failure is determined based on the ratio of the number of beats and the number of tones of several abnormal notes, wherein the abnormal notes are the recognized notes that do not match the training notes.

3. The method for comprehensive analysis of audio scores based on deep learning according to claim 2, characterized in that: The process of re-determining whether the recognition processing of the training audio information is qualified based on the total duration of the training audio information includes: The reason for the failure is determined based on the determination result, or the number of divided segments is corrected based on the total number of training notes.

4. The method for comprehensive analysis of audio scores based on deep learning according to claim 3, characterized in that: The number of divided segments is determined based on a comparison result of the total number of training notes with a preset total number of training notes, and the number of divided segments is in direct proportion to the total number of training notes.

5. The method for comprehensive analysis of audio scores based on deep learning according to claim 2, characterized in that: The process of determining the reason why the recognition processing of the training audio information is unqualified includes: Determining the reason for failure based on the comparison result of the number of beat mismatches and the number of pitch mismatches, and generating corresponding instructions based on the determined reason to determine the number of beats, pickup sensitivity or noise reduction ratio of the identified music score; The beat mismatch number is the number of abnormal notes whose recognized beats do not match the training beats, and the pitch mismatch number is the number of abnormal notes whose recognized pitches do not match the training pitches.

6. The method for comprehensive analysis of audio scores based on deep learning according to claim 5, characterized in that: The process of determining the number of beats of the identified music score based on the corresponding instruction includes: Calculating the beat difference between each training note and the corresponding abnormal note in sequence, and drawing a time-beat difference curve based on the beat difference; Calculating the slope variance based on the curve, and determining the number of beats of the identified music score information based on a comparison result of the slope variance with a preset slope variance, or redetermining the reason based on an average slope; Wherein, the slope of the curve corresponding to the abnormal note is obtained based on the curve, and the average slope is the average value between the slopes of each curve.

7. The method for comprehensive analysis of audio scores based on deep learning according to claim 6, characterized in that: The process of redetermining the cause based on the average slope includes: Based on the comparison result between the average slope and a preset average slope, it is determined whether to correct the beat number of the recognized music score information.

8. The method for comprehensive analysis of audio scores based on deep learning according to claim 7, characterized in that: When determining to correct the beat number of the identified music score information based on the average slope, the beat number of the identified music score information is increased or decreased based on a comparison result of the average slope with zero, and the amplitude of increase or decrease is proportional to the absolute value of the average slope.

9. The method for comprehensive analysis of audio scores based on deep learning according to claim 5, characterized in that: The process of determining the sound pickup sensitivity based on the corresponding instruction includes: Based on the comparison result of the abnormality ratio and the preset abnormality ratio, it is determined that the sound pickup sensitivity is increased, and the increase of the sound pickup sensitivity is proportional to the abnormality ratio; The abnormal proportion is the ratio between the number of abnormal notes and the total number of recognized notes.

10. The method for comprehensive analysis of audio scores based on deep learning according to claim 5, characterized in that: The process of determining the noise reduction ratio based on the corresponding instruction includes: Based on the comparison result of the abnormal tone ratio and the preset abnormal tone ratio, it is determined that the noise reduction ratio should be increased, and the increase of the noise reduction ratio is proportional to the abnormal tone ratio; The abnormal notes with mismatched pitch are recorded as abnormal pitch notes, and the abnormal pitch ratio is the ratio between the number of abnormal pitch notes and the total number of recognized notes.

Citation Information

Patent Citations

  • Audio information processing methods, devices and storage media

    CN108986841B