Method for recognizing music piece, computer device, and storage medium
Patent Information
- Application Number
- PCT/CN2026/083591
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-18
- Filing Date
- 2026-03-16
- Publication Date
- 2026-09-24
Smart Images

Figure CN2026083591_24092026_PF_FP_ABST
Abstract
Description
A method for identifying musical tracks, a computer device, and a storage medium.
[0001] Priority application
[0002] This application claims priority to Chinese patent applications filed on March 18, 2025, namely [2025103445151], "A method for identifying piano pieces, computer equipment and storage medium", and [2025103293996], "A method for identifying piano pieces, computer equipment and storage medium", which are incorporated herein by reference in their entirety. Technical Field
[0003] This application relates to the field of music information retrieval, and in particular to a method for identifying musical tracks, a computer device, and a storage medium. Background Technology
[0004] With the continuous advancement of music recognition technology and artificial intelligence, automatic audio recognition has become a key direction in the field of music information retrieval. Its core principle is to extract features from the audio (such as spectral peaks) to generate a unique audio fingerprint, and then efficiently match it with a pre-stored template library to achieve track identification.
[0005] For example, patent application CN111161758A discloses a method, system, and audio device for song recognition based on audio fingerprints. It collects song audio as template audio and obtains the corresponding spectrogram of the template audio. Peak points are extracted from the spectrogram as template audio fingerprints, and a template audio fingerprint database for song audio is constructed based on the template audio and the template audio fingerprints. The system also obtains the recorded audio of the current music and its corresponding spectrogram, extracting peak points from the spectrogram as recorded audio fingerprints. The recorded audio fingerprint is matched with the template audio fingerprints in the template audio fingerprint database. If the matching degree reaches a set threshold, the matching song audio is output, thus enabling automatic song recognition. The algorithm is efficient, accurate, and highly portable.
[0006] For example, patent application CN113486209A discloses an audio track recognition method, device, and readable storage medium. The method includes: acquiring the target audio to be identified, and after segmenting the target audio, extracting the audio fingerprint features of each segment; acquiring the song IDs that match each audio fingerprint feature; performing same-song clustering on the song IDs to obtain clustering results; using the clustering results to determine the confidence level of each song ID; and using the confidence level to select the target track from the tracks corresponding to each song ID.
[0007] However, piano playing involves complex chords formed by multiple notes sounding simultaneously, as well as rapid changes in speed and dynamics. Combined with the performer's personalized playing style, this can lead to distorted fingerprints, making it impossible to accurately match the correct music information and thus reducing the accuracy of music score recognition. Summary of the Invention
[0008] The main objective of this application is to provide a method for identifying musical tracks, a computer device, and a storage medium. To solve the aforementioned technical problems, this application specifically adopts the following technical solution:
[0009] A first aspect of this application is to provide a method for identifying piano pieces, the method comprising:
[0010] S101. Obtain the audio to be recognized;
[0011] S102. Analyze the audio data in the audio to be identified to identify several style transition nodes in the audio data, and divide the audio to be identified into several audio segments based on the several style transition nodes.
[0012] S103. Generate corresponding audio fingerprints based on each audio segment to obtain a plurality of first audio fingerprints;
[0013] S104. Determine the corresponding performance style features based on the audio data of each audio segment, and filter the general fingerprint database based on the performance style features to determine the target fingerprint database corresponding to each audio segment; wherein, the general fingerprint database includes multiple second audio fingerprints, and each second audio fingerprint is associated with performance style features and piano pieces;
[0014] S105. Compare the matching degree between the first audio fingerprint of each audio segment and the second audio fingerprint in the corresponding target fingerprint database to determine a number of candidate piano pieces corresponding to the audio segments; based on the matching degree and / or overlap of the number of candidate piano pieces, determine at least one target piano piece from the number of candidate piano pieces.
[0015] In some embodiments, the performance style characteristics include at least one of rhythmic characteristics, tempo characteristics, and genre characteristics.
[0016] In some embodiments, the method further includes: extracting contextual features of the audio data in the audio to be identified based on a preset model, and identifying several musical structures in the audio to be identified based on the contextual features; marking audio segments within each musical structure as continuous audio segments; and deleting the corresponding style transition node when the identified style transition node is located inside the continuous audio segment so that the continuous audio segment is not divided.
[0017] In some embodiments, after S102, the method further includes: obtaining the audio duration of each audio segment; when there are several consecutive audio segments whose audio duration is less than a preset interval threshold, and the total audio duration of the consecutive audio segments is greater than the preset interval threshold, merging the consecutive audio segments into one audio segment.
[0018] In some embodiments, the method further includes: when the total audio duration of consecutive audio segments is less than a preset interval threshold, classifying the corresponding audio segments as abnormal audio segments; or, when the audio duration of a single audio segment is less than a preset interval threshold, classifying the corresponding audio segment as an abnormal audio segment; comparing the first performance style feature of the preceding audio segment and the second performance style feature of the following audio segment of the abnormal audio segment; if the data difference between the first performance style feature and the second performance style feature is less than a preset difference threshold, merging the preceding audio segment, the abnormal audio segment, and the following audio segment into one audio segment.
[0019] In some embodiments, the method further includes: if a style transition node is not identified in the audio data, dividing the audio to be identified into several audio segments according to a preset segment duration, and / or dividing the audio to be identified into a preset number of audio segments.
[0020] In some embodiments, S105 includes: generating a corresponding audio fingerprint based on the audio to be identified to obtain a third audio fingerprint; comparing the third audio fingerprint of the audio to be identified with a second audio fingerprint in the general fingerprint database; when the matching degree between the second audio fingerprint and the first audio fingerprint is higher than a preset matching degree, determining the piano piece corresponding to the second audio fingerprint as a candidate piano piece.
[0021] In some embodiments, if the number of candidate piano pieces is less than a preset number, the method further includes: identifying the performance style characteristics of the audio to be identified based on the audio data of the audio to be identified; obtaining at least one representative piano piece based on the performance style characteristics of the audio to be identified; and / or obtaining at least one teaching piano piece based on the performance style characteristics of the audio to be identified; and using the representative piano piece and / or the teaching piano piece as the target piano piece.
[0022] A second aspect of this application is to provide a computer device, the device comprising:
[0023] Memory, used to store computer programs;
[0024] A processor is configured to execute the computer program and, in executing the computer program, implement the steps of the piano piece recognition method provided in any embodiment of this application.
[0025] A third aspect of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the piano piece recognition method provided in any embodiment of this application.
[0026] Beneficial technical effects:
[0027] This application proposes a method, computer device, and storage medium for identifying piano pieces. Stylistic pre-matching is performed on the audio and a fingerprint database based on typical and reliable characteristics in the audio. This ensures that the fingerprints used for subsequent comparisons share similar audio characteristics with the fingerprint database, thus adapting to dynamic performance scenarios involving personalized playing and cross-style adaptations in piano performance. Combined with a cross-validation matching mechanism, the accuracy of piano piece matching is improved. Furthermore, a simplified fingerprint database reduces the search scope and improves efficiency. Additionally, a music structure protection mechanism and a multi-style segment merging mechanism are introduced to avoid excessive audio fragmentation, enhancing fingerprint effectiveness and matching success rate, thereby improving the reliability and practicality of piano piece identification.
[0028] First, by detecting style transition nodes (such as rhythmic abrupt changes, tempo shifts, or genre transitions) in the audio, the performance audio is divided into multiple independent segments with a unified style. This ensures that each audio segment has more stable and less fluctuating characteristics under the same style, thereby generating more accurate and robust audio fingerprints. Furthermore, based on the performance style characteristics of the current segment, a target fingerprint database corresponding to the corresponding style is pre-screened to narrow the matching range, reduce the amount of interference data in the enriched general fingerprint database, and enhance the comparability of fingerprint features.
[0029] Finally, through the multi-segment cross-validation mechanism, even if a local segment fails to match due to performance adaptation or complex chords and dynamic changes, the target piano piece with higher credibility can still be determined by combining the majority of correct segments. This reduces the risk of misjudgment caused by the distortion of a single fingerprint. Furthermore, multiple segments can be matched in parallel, and the combination with the simplified target database balances matching efficiency and accuracy.
[0030] Furthermore, to ensure the accuracy of audio segmentation and the effectiveness of audio fingerprints, a music structure protection mechanism and a multi-style segment merging mechanism are introduced. For example, to ensure the integrity of the music structure and prevent key musical elements from being incorrectly segmented, when a style transition node is detected within the music structure, that node is automatically deleted to maintain the integrity of the music structure. Also, for example, if style transition nodes are misidentified due to fluctuations in performance details (such as temporary crescendos or ornaments) or environmental noise (such as audience coughs), resulting in overly fragmented audio segments, different strategies can be used to merge them, reducing the ambiguity caused by fragmented matching.
[0031] Furthermore, for performance audio without obvious stylistic changes or when it is difficult to identify a sufficient number of stylistic feature points, the audio is divided into equal segments of fixed duration or a predetermined number. When the number of candidate tracks for the target is insufficient, a style association recommendation mode is activated to recommend suitable teaching tracks or representative works based on the current performance characteristics, thereby enhancing user experience and practicality. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. The elements or parts in the drawings are not necessarily drawn to scale. Obviously, the drawings described below are some embodiments of this application; for those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0033] Figure 1 is a schematic flowchart of a piano repertoire recognition method provided in an embodiment of this application;
[0034] Figure 2 is a schematic diagram of the spectrum conversion of an audio to be identified according to an embodiment of this application;
[0035] Figure 3 is a schematic diagram of the spectral flux of an audio to be identified according to an embodiment of this application;
[0036] Figure 4 is a schematic diagram of the rhythmic features of an audio to be identified according to an embodiment of this application;
[0037] Figure 5 is a schematic block diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0039] In this document, suffixes such as “module,” “part,” or “unit” used to denote elements are used only for illustrative purposes and have no specific meaning in themselves. Therefore, “module,” “part,” or “unit” may be used interchangeably.
[0040] In this document, the terms "upper," "lower," "inner," "outer," "front," "rear," "one end," and "the other end," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0041] In this document, unless otherwise explicitly specified and limited, the terms "installed," "equipped with," and "connected," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0042] In this document, "and / or" includes any and all combinations of one or more of the listed related items.
[0043] In this article, "multiple" means two or more, that is, it includes two, three, four, five, etc.
[0044] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0045] This invention provides a method for music recognition, and the following will describe the solution in detail using piano music as an example:
[0046] In piano performance, each pianist has their own unique and personalized interpretation style. Subtle adjustments to rhythm, dynamics, and touch can cause the same piece to sound drastically different from one player's, resulting in variations in the audio fingerprint of the same piece. For example, some pianists might add subtle delays or accelerations to specific phrases, or alter the timbre through special touch techniques. These details often affect the final generated fingerprint, causing the audio characteristics originally intended to uniquely identify a musical segment to deviate, thus failing to accurately match the correct piece information. Furthermore, piano music contains various chord types (such as major triads, minor triads, augmented triads, etc.) that include not only the fundamental frequency but also various overtones, making misinterpretation more likely in complex chord cases.
[0047] Based on this, this application proposes a method, computer device, and storage medium for identifying piano pieces. Stylistic pre-matching is performed on the audio and fingerprint database based on typical and reliable characteristics in the audio. This ensures that the fingerprints used for subsequent comparisons share similar audio characteristics with the fingerprint database, thus adapting to dynamic performance scenarios such as personalized playing and cross-style adaptations in piano performance. Combined with a cross-validation matching mechanism, the accuracy of piano piece matching is improved. Simultaneously, a simplified fingerprint database reduces the search scope and improves efficiency. Furthermore, a music structure protection mechanism and a multi-style segment merging mechanism are introduced to avoid excessive audio fragmentation, improving the effectiveness of fingerprints and the success rate of matching, thereby enhancing the reliability and practicality of piano piece identification.
[0048] First, by detecting style transition nodes (such as rhythmic abrupt changes, tempo shifts, or genre transitions) in the audio, the performance audio is divided into multiple independent segments with a unified style. This ensures that each audio segment has more stable and less fluctuating characteristics under the same style, thereby generating more accurate and robust audio fingerprints. Furthermore, based on the performance style characteristics of the current segment, a target fingerprint database corresponding to the corresponding style is pre-screened to narrow the matching range, reduce the amount of interference data in the enriched general fingerprint database, and enhance the comparability of fingerprint features.
[0049] Finally, through the multi-segment cross-validation mechanism, even if a local segment fails to match due to performance adaptation or complex chords and dynamic changes, the target piano piece with higher credibility can still be determined by combining the majority of correct segments. This reduces the risk of misjudgment caused by the distortion of a single fingerprint. Furthermore, multiple segments can be matched in parallel, and the combination with the simplified target database balances matching efficiency and accuracy.
[0050] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other. Please refer to Figure 1, which is a schematic flowchart of a piano piece recognition method provided by an embodiment of this application. As shown in Figure 1, this application provides a piano piece recognition method, the method comprising steps S101 to S105.
[0051] S101. Obtain the audio to be recognized.
[0052] Specifically, an original audio file containing piano performance content can be collected from various possible sources as the audio to be identified. For example, an audio file recorded from a live performance can be used, or an audio file can be selected from existing audio files; there are no limitations on this.
[0053] S102. Analyze the audio data in the audio to be identified to identify several style transition nodes in the audio data, and divide the audio to be identified into several audio segments based on the several style transition nodes.
[0054] Audio data refers to the audio signal after the audio to be identified has been digitized, including but not limited to basic parameters such as sampling rate, bit depth, number of channels, and frequency domain characteristics, which are not limited here.
[0055] A style transition node refers to a point in time or a period of time within a continuous audio stream where the characteristics of music or sound undergo a significant change. These characteristics can include changes in pitch, rhythm, timbre, tempo, genre, etc., as well as shifts in emotional expression. It's important to note that since changes in audio characteristics may not occur instantaneously, a transition period is required. In this case, a style transition node can be a short period of time, such as 5 seconds.
[0056] For example, in a piano performance audio, the transition time between a fast, energetic section and a slow, lyrical section can be considered a style transition point. Similarly, the abrupt change from a steady beat to a free rhythm can also be seen as a style transition point.
[0057] It should be understood that some users prefer hybrid and improvisational performance styles, which are random and cannot be exhaustively listed, resulting in diverse audio data that is difficult to fully cover by simply expanding the size of a general fingerprint database. Therefore, generating audio fingerprints for audio segments with different audio features in the audio to be identified, and then pairing them with their corresponding target fingerprint databases, can effectively avoid interference between hybrid and improvisational audio data, improving the accuracy and comparability of fingerprint generation.
[0058] Specifically, the audio to be identified is analyzed to determine the corresponding audio data, from which various audio features are extracted from different dimensions, including but not limited to the dynamics, tempo, rhythm, and dynamic changes of the audio. By comparing the similarity between audio features in different time periods, time points with significant feature differences (such as abrupt rhythm changes, tempo shifts, or genre transitions) are identified as style transition nodes. When the style transition node is a time point, it is used as a segmentation point; when the style transition node is a time period, the start time of the time period can be used as the segmentation point. Based on these segmentation points, the audio to be identified is divided into several audio segments, each containing a relatively consistent musical style or performance style.
[0059] For example, if any two of the following are detected at time point A: a sudden change in rhythm, a change in tempo, or a change in genre, then the audio features at time point A are considered to have significant feature differences compared to audio features prior to time point A, and time point A is identified as a style transition node. Similarly, if the similarity between at least two audio features in time period B is lower than a preset similarity threshold, then the audio features in time period B are considered to have significant feature differences compared to audio features prior to time point B, and time period B is identified as a style transition node.
[0060] It should be understood that, based on the sensitivity requirements for identifying style transition nodes in practical applications, the criteria for identifying a certain time point as a style transition node can be flexibly set. For example, the number of audio features whose similarity is lower than a preset similarity threshold can be flexibly set to 1, 2, 3, etc.
[0061] In other words, different numbers of audio features can be used to define or determine whether a style shift has occurred.
[0062] For example, for style transition node detection, a large audio dataset labeled with various audio features and style transition nodes can be obtained. Based on this dataset, a classifier can be trained to automatically extract audio features from the audio to be identified and identify the style transition nodes contained therein. Alternatively, a deep learning model, such as a convolutional neural network, recurrent neural network, or long short-term memory network, can be trained on these audio datasets, and an optimization algorithm can be used to define the loss function. Thus, the trained deep learning model can automatically extract audio features from the audio to be identified and identify the style transition nodes contained therein.
[0063] S103. Generate corresponding audio fingerprints based on each audio segment to obtain several first audio fingerprints.
[0064] In this embodiment, the audio fingerprint is an identifier generated by extracting the unique digital features of the audio signal (such as an audio segment).
[0065] Specifically, an audio fingerprint is a compact digital feature extracted from an audio signal. It identifies the core features of an audio segment's content and can be represented as a sequence of numbers or a feature representation, without being limited here. Generating a corresponding audio fingerprint for each audio segment provides a foundation for subsequent matching, retrieval, and other operations.
[0066] For example, a Fast Fourier Transform (FFT) is used to convert the audio signal in the time domain to the frequency domain, thereby capturing the unique frequency patterns of each audio segment. Local amplitude maxima, i.e., frequency peaks, are found in the spectrogram, and the relationships between these peaks are encoded into a fingerprint using a hash function. This approach takes into account both the temporal relationships between individual frequency components and those between different frequency components, making the generated fingerprint more robust and unique.
[0067] For example, neural networks can be used to automatically extract audio features and generate more efficient audio fingerprints.
[0068] S104. Determine the corresponding performance style features based on the audio data of each audio segment, and filter the general fingerprint database based on the performance style features to determine the target fingerprint database corresponding to each audio segment; wherein, the general fingerprint database includes multiple second audio fingerprints, and each second audio fingerprint is associated with performance style features and piano pieces.
[0069] Among them, performance style features refer to the set of feature parameters obtained by quantitatively analyzing audio data from dimensions such as dynamics, rhythm, tempo, and genre. It should be understood that when a musical performance enters a new section, or changes in performance style or performer, various audio features may undergo significant changes. The specific moment of transition from one type of audio feature to another is called a style transition node, that is, a turning point within the music. A subset of various audio features is selected as performance style features to describe the consistent characteristics of the entire audio segment from a specific perspective. After dividing the audio to be identified into multiple audio segments based on style transition nodes, the performance style features within each audio segment are relatively uniform. Therefore, multiple audio segments possess more stable and less fluctuating audio features, resulting in a more robust audio fingerprint. Simultaneously, the consistent characteristics of the audio are used to filter the target fingerprint database to narrow the matching range. In other words, audio features collaboratively construct a hierarchical music description system from multiple dimensions. These features from different dimensions can be used for audio segmentation at style transition nodes and to filter the target fingerprint database to narrow the matching range.
[0070] In some embodiments, the performance style characteristics include at least one of rhythmic characteristics, tempo characteristics, and genre characteristics. Rhythmic characteristics describe the temporal sequence of note value combinations and beat strength distribution, such as beat cycle, rhythm type, and note density. Tempo characteristics describe the temporal sequence of the music's speed, such as overall tempo and local tempo changes (e.g., crescendo, crescendo, free tempo). Genre characteristics describe the style category to which the music belongs, reflecting compositional techniques and performance conventions, such as Baroque, Classical, Romantic, and Modern styles.
[0071] It should be understood that rhythmic features, tempo features, and genre features have a certain degree of complementarity. Rhythmic and tempo features can address the issue of personalized adaptations by performers during performances. Even if a performer speeds up or slows down certain pieces, the audio fingerprint can still be matched with the corresponding tempo feature version for comparison. Genre features resolve cross-style mismatches. Even if a performer switches between various styles and genres during a performance, the audio fingerprint can still be matched with the corresponding genre feature version for comparison. This refines the fingerprint database and improves the accuracy and speed of comparison.
[0072] For example, an automatic beat tracking algorithm can be used to identify rhythmic patterns and tempo changes in audio clips, thereby determining rhythmic and tempo characteristics, as well as genre characteristics. It should be understood that different performance styles are often accompanied by specific rhythmic structures and dynamic variations; for instance, classical music tends to have more complex rhythms and slower tempos, while jazz may include improvisation and rapid rhythmic changes.
[0073] For example, melody extraction and chord recognition algorithms are used to analyze the melodic line and harmonic progression of audio clips, thereby determining genre characteristics. It should be understood that different styles of music differ significantly in melody construction and harmonic selection, such as the simple and repetitive melodic structures common in popular music, and the complex harmonic progressions in classical music.
[0074] The universal fingerprint database is a comprehensive and rich audio fingerprint database containing a large number of audio fingerprints for piano pieces. These audio fingerprints are unique identifiers generated by feature extraction from the original audio data, which includes different versions of piano pieces performed in different styles. Each audio fingerprint is associated with corresponding identity data, such as performance style characteristics and information about the musical work (composer, piece title), providing a wealth of reference for subsequent screening.
[0075] Specifically, based on known performance style characteristics, such as specific tempo, rhythm, and genre, corresponding filtering conditions are set. These filtering conditions are used to screen the general fingerprint database. For example, if an audio segment is detected to have obvious jazz style characteristics, then the second audio fingerprint with jazz style characteristics can be included in the target fingerprint database. It should be understood that the target fingerprint database is a subset selected from the general fingerprint database and constructed for the specific performance style characteristics of each audio segment. The target fingerprint database contains the second audio fingerprint that best matches the input audio segment in terms of performance style characteristics.
[0076] In some embodiments, before filtering the general fingerprint database based on the performance style features, the method further includes: for the extracted candidate features, filtering the candidate features based on the confidence level of each candidate feature, and selecting candidate features with a confidence level higher than a preset confidence level as performance style features. This uses more typical and reliable performance style features from the audio clips to set corresponding filtering conditions, thereby effectively filtering the target fingerprint database while avoiding over-filtering.
[0077] It should be understood that, in order to provide rich and sufficient comparison samples in the subsequent matching process, the general fingerprint database includes audio fingerprints of a large number of different performance versions of piano pieces. For example, there are multiple performance versions of the same piano piece under different tempo features, different rhythm features, and different genre features. This results in a certain amount of data redundancy in the general fingerprint database. However, performance style features can be used to pre-screen these comparison samples to obtain a highly relevant and customized subset fingerprint database (i.e., the target fingerprint database), which speeds up the subsequent retrieval and improves the relevance of the retrieval results.
[0078] In some embodiments, the detection of style transition nodes and the analysis of performance style features both rely on specific parameters in the audio data, such as spectral characteristics and rhythmic patterns. This can be achieved by using a pre-trained machine model as described in the previous embodiments, or by using a spectrogram feature analysis algorithm.
[0079] In some embodiments, a spectrogram of the audio data corresponding to the audio to be identified is generated based on the Fast Fourier Transform (FFT). A spectrogram is a visualization tool representing frequency components that change over time, showing the frequency distribution and energy intensity of the audio signal at different points in time. Spectral characteristics of the audio, such as spectral centroid, spectral flux, and root-mean-square energy, are calculated from the spectrogram. These characteristics reflect the timbre, rhythm, and dynamic changes of the audio. Furthermore, the audio is segmented using a sliding window technique, and the feature differences between adjacent segments are compared. If multiple features begin to change significantly at a certain node, that node may be a style transition node. Alternatively, the performance style characteristics of each audio segment can be analyzed based on the spectral characteristics of the spectrogram of each audio segment.
[0080] Taking the analysis of the rhythmic features of an audio segment based on spectrogram technology as an example, please refer to Figures 2 to 4. Figure 2 is a schematic diagram of the spectrum transformation of an audio to be identified provided by an embodiment of this application; Figure 3 is a schematic diagram of the spectral flux of an audio to be identified provided by an embodiment of this application; Figure 4 is a schematic diagram of the rhythmic features of an audio to be identified provided by an embodiment of this application.
[0081] As shown in Figure 2, the upper figure is an audio waveform diagram of a portion of the audio to be identified, displaying the amplitude of the sound signal changing over time. This is used to understand the basic rhythm and volume changes of the audio; for example, observing periodic changes in the audio signal may represent repeated notes or melodies in a piano piece. The lower figure is a Mel spectrogram diagram of a portion of the audio to be identified, using color changes to show the energy distribution of different frequency components, used for more in-depth analysis of the audio content.
[0082] Furthermore, a spectral flux map is generated based on the Mel spectrogram, which reflects the changes in frequency components of the audio signal over time. This is used to identify the beginning and end of notes and detect rhythmic changes. For example, frequently occurring tall bars may indicate rapid rhythmic changes, while more uniform bars may indicate a more stable rhythm. The spectral flux map shown in Figure 3 presents the spectral flux of a portion of the audio to be identified.
[0083] Furthermore, the rhythmic features of the audio to be identified are analyzed based on the spectral flux plot. As shown in Figure 4, the upper figure is a beat estimation plot based on spectral flux. The blue bars represent the spectral flux value at each time point, and the green dashed line marks the beat position predicted by existing beat tracking algorithms. The lower figure is a cumulative score plot. The cumulative score is a score calculated based on spectral flux and other features, used to evaluate whether a certain time point may be a beat point. The orange curve represents the cumulative score, and the green dashed line also marks the estimated beat position. The cumulative score can more accurately determine the beat position in the audio. Therefore, based on the beat position, the rhythmic features in the audio to be identified can be analyzed and determined. These rhythmic features can be used for the detection of style transition nodes and the analysis of performance style features.
[0084] S105. Compare the matching degree between the first audio fingerprint of each audio segment and the second audio fingerprint in the corresponding target fingerprint database to determine a number of candidate piano pieces corresponding to the audio segments; based on the matching degree and / or overlap of the number of candidate piano pieces, determine at least one target piano piece from the number of candidate piano pieces.
[0085] Specifically, for any first audio fingerprint, it is compared with the corresponding second audio fingerprint in the target fingerprint database. The comparison process can employ efficient algorithms, such as hash tables or inverted indexes, to achieve fast and accurate matching. For each matching result, a matching degree is calculated based on the similarity between the two audio fingerprints; the higher the matching degree, the more likely the two audio segments come from the same piano piece. At least one audio fingerprint with a matching degree higher than a preset matching degree and / or the highest matching degree is selected as the target audio fingerprint. The piano pieces associated with the target audio fingerprint are then obtained as candidate piano pieces. Based on the matching degree of these candidate piano pieces and their potential overlap (i.e., the frequency with which multiple audio segments point to the same piece), at least one target piano piece that best matches the input audio segment can be selected from numerous candidate pieces. This allows for quick and accurate identification of the music the user is listening to or wants to learn about, improving the user experience.
[0086] Therefore, through the multi-segment cross-validation mechanism, even if a local segment fails to match due to performance adaptation or complex chords and dynamic changes, these interferences can be eliminated by combining the two dimensions of matching degree and overlap. For example, by combining the majority of correct segments, a more credible target piano piece can be determined, reducing the risk of misjudgment caused by the distortion of a single fingerprint. Furthermore, multiple segments can be matched in parallel, and the combination with a simplified target database balances matching efficiency and accuracy.
[0087] Furthermore, due to interference from factors such as fluctuations in performance details (e.g., temporary crescendos and ornamentation) and environmental noise (e.g., audience coughing), the identification of style transition nodes may be misrecognized, leading to over-segmentation of audio segments. To ensure the accuracy of audio segment segmentation and the effectiveness of audio fingerprints, a music structure protection mechanism is introduced.
[0088] In some embodiments, to ensure the integrity of the music structure and prevent key music elements from being incorrectly segmented, the method further includes: extracting contextual features of the audio data in the audio to be identified based on a preset model, and identifying several music structures in the audio to be identified based on the contextual features; marking audio segments within each music structure as continuous audio segments; and deleting the corresponding style transition node when the identified style transition node is located inside the continuous audio segment so that the continuous audio segment is not segmented.
[0089] Contextual features refer to the set of high-order features extracted from audio that reflect the semantic coherence of the music. For example, structural context refers to the organizational form of musical passages or phrases, such as repetition, variation, and development. By comprehensively utilizing these contextual features, we can avoid treating notes in isolation and understand the inherent connections within the music.
[0090] Musical structure refers to the hierarchical units within a musical work that follow formal logic, including different levels of structural forms such as measures, phrases, and recurring themes. A measure is the basic unit of time in a musical score, composed of a certain number of beats, representing a complete rhythmic cycle. A phrase is a musical expression unit consisting of several measures, typically containing a complete thought or emotional expression. A recurring theme refers to a musical segment that appears repeatedly in a musical work.
[0091] Specifically, the system extracts contextual features from the audio to be identified using a pre-defined model to analyze and identify musical structural boundaries (such as phrases, sections, or recurring themes). Continuous audio segments within the same structure are marked as indivisible wholes, i.e., continuous audio segments. When a style transition node is located within a continuous audio segment, it is automatically determined that the node is local interference rather than a genuine style change, and thus deleted to avoid disrupting the semantic units of the music. For example, some ornaments may appear in the middle of a phrase; in this case, the ornament might be identified as a style transition node. To maintain the integrity and coherence of each audio segment, the style transition node is deleted, avoiding information loss or misinterpretation due to misjudgment.
[0092] It should be understood that context awareness avoids over-segmentation caused by fluctuations in performance details (such as temporary crescendos and ornamentation) and noise, ensuring the integrity of key musical motifs and the contextual relevance of fingerprint features, thus improving the accuracy and efficiency of audio recognition. Especially when dealing with complex and varied musical works, it can effectively suppress misjudgments caused by improvisation, for example, complex rhythmic changes in a jazz improvisational cadenza are still classified within the same improvisational structural unit.
[0093] In some embodiments, a large-scale dataset with labeled musical structures is used, containing piano pieces from genres such as classical, jazz, and pop. The annotation information covers musical structure boundaries (e.g., phrases, sections, repetitions), style labels (e.g., Baroque, Romantic), and performance events (ornaments, cadenzas). This dataset is used to train a base model under the Transformer architecture. Long-range contextual dependencies are modeled using a self-attention mechanism to parse macroscopic structural logic such as melody development and harmonic progression, resulting in the pre-defined model. It should be noted that the pre-defined model can also use other model architectures or other types of datasets; this is not limited here.
[0094] Furthermore, for excessively fragmented audio segments, a multi-style segment merging mechanism is introduced. By merging invalid short segments, the number of fingerprint generation and matching operations is reduced, avoiding mismatches caused by fragmented fingerprints while maintaining the correct segmentation of other audio segments. It should be understood that excessively short audio segments may lead to insufficient feature extraction, resulting in a high matching degree between the generated first audio fingerprint and a large number of fingerprints in the fingerprint database, thus filtering out a large number of distracting alternative piano pieces. Targeted processing of these fragmented audio segments reduces the amount of data required for subsequent matching operations and avoids the error accumulation problem caused by excessively short segments, making subsequent audio recognition more accurate and reliable.
[0095] In some embodiments, after S102, the method further includes: when there are several consecutive audio segments whose audio duration is less than a preset interval threshold, and the total audio duration of the consecutive audio segments is greater than the preset interval threshold, merging the consecutive audio segments into one audio segment.
[0096] The preset interval threshold is used to determine whether the length of a single audio segment contains enough information for effective analysis or processing. Specific values can be flexibly set according to the actual application scenario and are not limited here.
[0097] Specifically, when the individual audio durations of multiple consecutive audio segments are less than a preset interval threshold, but their total duration reaches the preset interval threshold, these audio segments are merged into one audio segment. At this time, the existence of multiple fragmented audio segments may correspond to characteristic performance segments in piano playing, which are designed with rapid changes in rhythm, speed, style, etc., thus leading to them being divided into multiple overly short audio segments. In order to accurately capture the identifiable features (i.e., rapid changes) in this audio, multiple fragmented audio segments are merged, retaining the effective audio information. Subsequently, only one audio fingerprint is generated. Even if these fragmented audio segments are invalid audio, the number of interfering alternative piano pieces generated is small, and the impact on the final recognition result is controllable.
[0098] In some embodiments, the audio duration of each audio segment is obtained; when the audio duration of a single audio segment is less than a preset interval threshold, the corresponding audio segment is merged with the adjacent audio segments into one audio segment.
[0099] Specifically, when both the preceding and following audio segments of an audio segment are longer than a preset interval threshold, and the audio duration of the audio segment is shorter than the preset interval threshold, the corresponding audio segment is a single audio segment. This audio segment will be merged with the adjacent segments to ensure that each segment reaches a reasonable length.
[0100] In some embodiments, the method further includes: when the total audio duration of consecutive audio segments is less than a preset interval threshold, classifying the corresponding audio segments as abnormal audio segments; or, when the audio duration of a single audio segment is less than a preset interval threshold, classifying the corresponding audio segment as an abnormal audio segment; comparing the first performance style feature of the preceding audio segment and the second performance style feature of the following audio segment of the abnormal audio segment; if the data difference between the first performance style feature and the second performance style feature is less than a preset difference threshold, merging the preceding audio segment, the abnormal audio segment, and the following audio segment into one audio segment.
[0101] Specifically, when an abnormal audio segment is detected (the duration of a single segment or the total duration of consecutive segments is less than a preset threshold), the first and second performance style features of its preceding and following audio segments (i.e., the preceding and following audio segments) are extracted, and style consistency is determined by calculating the feature difference value. If the data difference between the preceding and following audio segments is less than a preset difference threshold, it indicates that the preceding and following segments of the abnormal audio segment have consistency or similarity in musical style. Therefore, the abnormal audio segment is highly likely to be a misclassification, and these three audio segments are merged into a longer audio segment.
[0102] For example, the preset difference threshold can be flexibly set according to different performance style characteristics to verify whether two audio clips contain relatively consistent musical styles or performance forms. For rhythmic features, the preset difference threshold can be set based on the similarity of rhythm types between two rhythmic features, such as a preset difference threshold of 80% similarity; for tempo features, the preset difference threshold can be set based on the difference in beats per minute (BPM) between two tempo features, such as a preset difference threshold of 15 BPM; for example, for genre features, the preset difference threshold can be set based on whether two genre features are the same, such as a preset difference threshold of 100% genre feature matching.
[0103] In some embodiments, when the total audio duration of consecutive audio segments is less than a preset interval threshold, the corresponding audio segments are designated as abnormal audio segments; or, when the audio duration of a single audio segment is less than a preset interval threshold, the corresponding audio segment is designated as an abnormal audio segment; the first performance style feature of the preceding audio segment and the second performance style feature of the following audio segment are compared; if the data difference between the first performance style feature and the second performance style feature is greater than a preset difference threshold, the abnormal audio segment is merged with the preceding or following audio segment into one audio segment.
[0104] In some embodiments, when the total audio duration of consecutive audio segments is less than a preset interval threshold, the corresponding audio segments are identified as abnormal audio segments. The first performance style feature of the preceding audio segment and the second performance style feature of the following audio segment are compared. If the data difference between the first performance style feature and the second performance style feature is greater than a preset difference threshold, the abnormal audio segment is removed and not included in steps S103 to S105. This avoids the first audio fingerprint generated from multiple short consecutive audio segments from matching a large number of interfering backup piano pieces, thereby preventing a decrease in the accuracy of audio recognition.
[0105] In some embodiments, the method further includes: if a style transition node is not identified in the audio data, dividing the audio to be identified into several audio segments according to a preset segment duration, and / or dividing the audio to be identified into a preset number of audio segments.
[0106] Specifically, if a clear style transition node cannot be identified, the audio to be identified can be divided based on a preset segment length or a preset number of segments. The entire audio to be identified can be divided into multiple parts according to the set time length; alternatively, the entire audio to be identified can be evenly distributed according to the required total number of segments, ensuring that each segment has approximately the same length. It should be understood that when the style transition node is not obvious, the style uniformity of the entire audio to be identified can be understood, and the audio segments can be divided using a simple and effective segmentation standard to achieve a multi-segment cross-validation mechanism.
[0107] In some embodiments, S105 includes: generating a corresponding audio fingerprint based on the audio to be identified to obtain a third audio fingerprint; comparing the third audio fingerprint of the audio to be identified with a second audio fingerprint in the general fingerprint database; when the matching degree between the second audio fingerprint and the first audio fingerprint is higher than a preset matching degree, determining the piano piece corresponding to the second audio fingerprint as a candidate piano piece.
[0108] Specifically, based on generating the first audio fingerprint from audio segments, an additional global third audio fingerprint is generated for the complete audio to be identified. Through a two-layer verification mechanism of local segment fingerprints and global fingerprints, the corresponding music piece for the audio to be identified is accurately identified. The global fingerprint (third audio fingerprint) of the audio to be identified is compared with the second audio fingerprint in the general fingerprint database. When the matching degree is higher than the preset matching degree, it is determined as a candidate piano piece.
[0109] It should be understood that segment fingerprints (first audio fingerprints) enhance adaptability to performance adaptations and local variations by focusing on the detailed features of stylized segments (such as rhythmic patterns and combinations of ornaments), and are particularly suitable for identifying mixed styles or improvisational passages; while overall fingerprints (third audio fingerprints) capture the macroscopic structural features of the audio (such as musical form and thematic repetition patterns), making up for the cross-segment contextual relationships that may be lost in segment fingerprints.
[0110] From another perspective, this invention actually provides a hierarchical matching scheme with matching precision ranging from high to low. Specifically, this invention first uses smaller-scale audio segments for initial large-scale identification (using smaller audio segments for refined identification, and initially filtering out multiple matching results through multiple audio segments). Subsequently, this invention uses larger-scale audio data to quickly filter the refined matching results (e.g., filtering directly based on the audio to be identified). This hierarchical matching scheme with precision ranging from high to low can reduce the probability of missing or incorrectly selecting tracks through refined matching, while quickly eliciting the final target from multiple candidates through large-scale rapid filtering.
[0111] In some embodiments, if the number of candidate piano pieces is less than a preset number, the method further includes: identifying the performance style characteristics of the audio to be identified based on the audio data of the audio to be identified; obtaining at least one representative piano piece based on the performance style characteristics of the audio to be identified; and / or obtaining at least one teaching piano piece based on the performance style characteristics of the audio to be identified; and using the representative piano piece and / or the teaching piano piece as the target piano piece.
[0112] Specifically, when the number of candidate piano pieces is lower than the preset number, the number of matching candidate piano pieces is insufficient. An association recommendation mode is triggered based on performance style characteristics to enhance the practicality of audio recognition. The data of the audio to be recognized is analyzed to identify its performance style characteristics, and based on this, at least one representative piano piece or teaching piano piece that matches it is obtained.
[0113] For music appreciation scenarios, recommended representative pieces help users discover classic works of the same style. These recommendations are sorted by style similarity, selecting authoritative and well-known classic works within that style. For example, when Baroque ornamentation features are identified, excerpts from "The Well-Tempered Clavier" are recommended. For music learning scenarios, recommended teaching pieces provide targeted practice guidance. Therefore, teaching piece recommendations are related to the needs of the teaching scenario, matching practice pieces with appropriate technical difficulty. For example, when classical genre characteristics are detected, "Für Elise" and "Turkish March" are recommended.
[0114] In some embodiments, a style repertoire mapping map is pre-constructed, which records the mapping relationship between different performance style features and various piano pieces. Then, after extracting the performance style features such as rhythm, tempo, and genre of the audio to be identified, it can be mapped to the style repertoire mapping map to output the corresponding representative piano pieces or teaching piano pieces.
[0115] In some embodiments, the user's audio recognition purpose is collected in advance to distinguish between music appreciation scenarios and music learning scenarios, thereby accurately pushing at least one representative piano piece or teaching piano piece according to the user's needs.
[0116] In some embodiments, the performance style features also include dynamic features, which quantify the amplitude and transition patterns of dynamic contrasts (such as crescendo and decrescendo) in the performance by analyzing the energy distribution of the audio to be identified in the spectrogram.
[0117] It should be noted that this invention actually provides a method for music score recognition. The above application of piano pieces is merely an illustrative example to illustrate the music score recognition method. This invention can also be applied to other types of instrumental performances, such as recordings for string instruments (violin, cello), wind instruments (flute, saxophone), plucked instruments (guitar, guzheng), and other types of instruments. This invention does not limit these applications.
[0118] For example, the music score recognition method provided by this invention includes:
[0119] S101. Obtain the audio to be recognized;
[0120] S102. Analyze the audio data in the audio to be identified to identify several style transition nodes in the audio data, and divide the audio to be identified into several audio segments based on the several style transition nodes.
[0121] S103. Generate corresponding audio fingerprints based on each audio segment to obtain a plurality of first audio fingerprints;
[0122] S104. Determine the corresponding performance style features based on the audio data of each audio segment, and filter the general fingerprint database based on the performance style features to determine the target fingerprint database corresponding to each audio segment; wherein, the general fingerprint database includes multiple second audio fingerprints, and each second audio fingerprint is associated with performance style features and a song;
[0123] S105. Compare the matching degree between the first audio fingerprint of each audio segment and the second audio fingerprint in the corresponding target fingerprint database to determine a number of candidate tracks corresponding to the audio segments; based on the matching degree and / or overlap of the number of candidate tracks, determine at least one target track from the number of candidate tracks.
[0124] Correspondingly, the present invention also provides a music score recognition system, which includes:
[0125] The recognition module is used to acquire the audio to be recognized;
[0126] The analysis module is used to analyze the audio data in the audio to be identified, to identify several style transition nodes in the audio data, and to divide the audio to be identified into several audio segments based on the several style transition nodes.
[0127] The fingerprint generation module is used to generate corresponding audio fingerprints based on each audio segment to obtain a plurality of first audio fingerprints;
[0128] The first filtering module is used to determine the corresponding performance style features based on the audio data of each audio segment, and to filter the general fingerprint database based on the performance style features to determine the target fingerprint database corresponding to each audio segment; wherein, the general fingerprint database includes multiple second audio fingerprints, and each second audio fingerprint is associated with performance style features and a piece of music;
[0129] The second filtering module is used to compare the matching degree between the first audio fingerprint of each audio segment and the second audio fingerprint in the corresponding target fingerprint database to determine the candidate tracks corresponding to several audio segments; and to determine at least one target track from the several candidate tracks based on the matching degree and / or overlap degree of the several candidate tracks.
[0130] For example, the above method can be implemented as a computer program that can run on the computer device shown in FIG5.
[0131] As shown in Figure 5, the computer device includes a processor, memory, and network interface connected via a system bus. The memory may include non-volatile storage media and internal memory. The non-volatile storage media may store the operating system and computer programs. The computer programs include program instructions that, when executed, cause the processor to perform a method for recognizing any piano piece.
[0132] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0133] Internal memory provides an environment for the execution of computer programs stored in non-volatile storage media. When executed by a processor, the computer program enables the processor to perform any method for recognizing piano pieces.
[0134] This network interface is used for network communication, such as sending assigned tasks.
[0135] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0136] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps:
[0137] S101. Obtain the audio to be recognized;
[0138] S102. Analyze the audio data in the audio to be identified to identify several style transition nodes in the audio data, and divide the audio to be identified into several audio segments based on the several style transition nodes.
[0139] S103. Generate corresponding audio fingerprints based on each audio segment to obtain a plurality of first audio fingerprints;
[0140] S104. Determine the corresponding performance style features based on the audio data of each audio segment, and filter the general fingerprint database based on the performance style features to determine the target fingerprint database corresponding to each audio segment; wherein, the general fingerprint database includes multiple second audio fingerprints, and each second audio fingerprint is associated with performance style features and piano pieces;
[0141] S105. Compare the matching degree between the first audio fingerprint of each audio segment and the second audio fingerprint in the corresponding target fingerprint database to determine a number of candidate piano pieces corresponding to the audio segments; based on the matching degree and / or overlap of the number of candidate piano pieces, determine at least one target piano piece from the number of candidate piano pieces.
[0142] For example, the processor is used to run a computer program stored in the memory, and is also used to implement the steps of the piano piece recognition method provided in any embodiment of this application, which will not be repeated here.
[0143] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement the steps of the piano piece recognition method provided in any of the embodiments of this application.
[0144] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0145] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for identifying piano pieces, characterized in that, The method includes: S101. Obtain the audio to be recognized; S102. Analyze the audio data in the audio to be identified to identify several style transition nodes in the audio data, and divide the audio to be identified into several audio segments based on the several style transition nodes. S103. Generate corresponding audio fingerprints based on each audio segment to obtain a plurality of first audio fingerprints; S104. Determine the corresponding performance style features based on the audio data of each audio segment, and filter the general fingerprint database based on the performance style features to determine the target fingerprint database corresponding to each audio segment; wherein, the general fingerprint database includes multiple second audio fingerprints, and each second audio fingerprint is associated with performance style features and piano pieces; S105. Compare the matching degree between the first audio fingerprint of each audio segment and the second audio fingerprint in the corresponding target fingerprint database to determine a number of candidate piano pieces corresponding to the audio segments; based on the matching degree and / or overlap of the number of candidate piano pieces, determine at least one target piano piece from the number of candidate piano pieces.
2. The method according to claim 1, characterized in that, The performance style characteristics include at least one of rhythmic characteristics, tempo characteristics, and genre characteristics.
3. The method according to claim 1, characterized in that, The method further includes: The contextual features of the audio data in the audio to be identified are extracted based on a preset model, and several musical structures in the audio to be identified are identified based on the contextual features. Each audio segment within the aforementioned music structure is marked as a consecutive audio segment; When the identified style transition node is located inside the continuous audio segment, the corresponding style transition node is deleted so that the continuous audio segment is not divided.
4. The method according to claim 1, characterized in that, Following S102, the following is also included: Obtain the audio duration of each audio segment; When there are several consecutive audio segments whose audio duration is less than a preset interval threshold, and the total audio duration of the consecutive audio segments is greater than the preset interval threshold, the consecutive audio segments are merged into one audio segment.
5. The method according to claim 4, characterized in that, The method further includes: When the total audio duration of consecutive audio segments is less than a preset interval threshold, the corresponding audio segments are considered abnormal audio segments; or, when the audio duration of a single audio segment is less than a preset interval threshold, the corresponding audio segment is considered abnormal audio segment. Compare the first performance style characteristics of the preceding audio segment and the second performance style characteristics of the subsequent audio segment of the abnormal audio segment; If the data difference between the first performance style feature and the second performance style feature is less than a preset difference threshold, the preceding audio segment, the abnormal audio segment, and the subsequent audio segment are merged into one audio segment.
6. The method according to claim 1, characterized in that, The method further includes: If the style transition node in the audio data is not identified, the audio to be identified is divided into several audio segments according to the preset segment duration, and / or the audio to be identified is divided into a preset number of audio segments.
7. The method according to claim 1 or 6, characterized in that, S105 includes: Based on the audio to be identified, a corresponding audio fingerprint is generated to obtain a third audio fingerprint; The third audio fingerprint of the audio to be identified is compared with the second audio fingerprint in the general fingerprint database. When the matching degree between the second audio fingerprint and the first audio fingerprint is higher than the preset matching degree, the piano piece corresponding to the second audio fingerprint is determined as the candidate piano piece.
8. The method according to claim 1, characterized in that, If the number of candidate piano pieces is less than a preset number, the method further includes: Identify the performance style characteristics of the audio to be identified based on the audio data of the audio to be identified; Based on the performance style characteristics of the audio to be identified, at least one representative piano piece is obtained; and / or, based on the performance style characteristics of the audio to be identified, at least one teaching piano piece is obtained. The representative piano pieces and / or teaching piano pieces are used as the target piano pieces.
9. A computer device, characterized in that, The device includes: Memory, used to store computer programs; A processor is configured to execute the computer program and, in executing the computer program, implement the method for recognizing piano pieces as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to implement the method for recognizing piano pieces as described in any one of claims 1 to 8.