Performance analysis method, electronic device, readable storage medium, and program product

By identifying and filtering the note feature information of the musical instrument performance and matching it with the beat template information of the target content according to the beat, the problem of inaccurate performance analysis in the prior art is solved, and higher analysis accuracy and reliability are achieved.

CN120032609APending Publication Date: 2025-05-23TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133841.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately analyze the performance contents performed by musical instruments, resulting in inaccurate performance analysis.

Method used

By obtaining the audio spectrum information of the instrument's performance, identifying multiple note feature information, filtering the initial note event to obtain the target note event, and dividing it according to the beat, matching the beat template information with the target content to obtain the analysis results.

Benefits of technology

It improves the reliability and accuracy of note events, ensures the reliability and accuracy of beat information to be analyzed, and thus improves the reliability and accuracy of performance analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032609A_ABST
    Figure CN120032609A_ABST
Patent Text Reader

Abstract

The invention relates to a playing analysis method, electronic equipment, a computer readable storage medium and a computer program product, and relates to the technical field of audio processing and artificial intelligence. The reliability and accuracy of playing analysis can be improved. The method comprises the following steps: acquiring various note feature information obtained by identifying audio frequency spectrum information of target content, and obtaining an initial note event obtained by playing the target content by a musical instrument according to the various note feature information, filtering the initial note event according to the note duration of the notes contained in the initial note event to obtain a target note event obtained by the musical instrument playing target content, and dividing the notes contained in the target note event according to beats to obtain to-be-analyzed beat information, and obtaining an analysis result for musical instrument playing according to a matching result of the beat information to be analyzed and the beat template information of the target content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing and artificial intelligence, and in particular to a performance analysis method, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the development of artificial intelligence technology, artificial intelligence technology has been applied to various fields and scenarios. In the field of audio, artificial intelligence technology can be applied to the analysis of audio data. For example, in music applications, artificial intelligence technology can be used to analyze the user's vocal singing songs and analyze the degree of matching with the original singer. In the current technology, the method of analyzing vocal singing songs is difficult to apply to the performance analysis of related content using musical instruments, and there is a technical problem of inaccurate performance analysis. Summary of the invention

[0003] Based on this, it is necessary to provide a performance analysis method, an electronic device, a computer-readable storage medium and a computer program product to address the above technical problems.

[0004] In a first aspect, the present application provides a performance analysis method, comprising:

[0005] Acquire multiple note feature information obtained by identifying audio spectrum information of target content; the audio spectrum information is audio spectrum information obtained by playing the target content with a musical instrument;

[0006] According to the multiple note feature information, an initial note event obtained by playing the target content with a musical instrument is obtained; the initial note event includes multiple notes;

[0007] Filtering the initial note event according to the note duration of the note contained in the initial note event to obtain a target note event obtained by a musical instrument playing the target content;

[0008] Dividing the notes contained in the target note event according to beats to obtain beat information to be analyzed;

[0009] According to the matching result of the beat information to be analyzed and the beat template information of the target content, an analysis result for the instrument performance is obtained.

[0010] In a second aspect, the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0011] Acquire multiple note feature information obtained by identifying audio spectrum information of target content; the audio spectrum information is audio spectrum information obtained by a musical instrument playing the target content; based on the multiple note feature information, obtain an initial note event obtained by a musical instrument playing the target content; the initial note event contains multiple notes; filter the initial note event based on the note duration of the notes contained in the initial note event to obtain a target note event obtained by a musical instrument playing the target content; divide the notes contained in the target note event according to beats to obtain beat information to be analyzed; and obtain an analysis result for the musical instrument performance based on a matching result between the beat information to be analyzed and the beat template information of the target content.

[0012] In a third aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0013] Acquire multiple note feature information obtained by identifying audio spectrum information of target content; the audio spectrum information is audio spectrum information obtained by a musical instrument playing the target content; based on the multiple note feature information, obtain an initial note event obtained by a musical instrument playing the target content; the initial note event contains multiple notes; filter the initial note event based on the note duration of the notes contained in the initial note event to obtain a target note event obtained by a musical instrument playing the target content; divide the notes contained in the target note event according to beats to obtain beat information to be analyzed; and obtain an analysis result for the musical instrument performance based on a matching result between the beat information to be analyzed and the beat template information of the target content.

[0014] In a fourth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0015] Acquire multiple note feature information obtained by identifying audio spectrum information of target content; the audio spectrum information is audio spectrum information obtained by a musical instrument playing the target content; based on the multiple note feature information, obtain an initial note event obtained by a musical instrument playing the target content; the initial note event contains multiple notes; filter the initial note event based on the note duration of the notes contained in the initial note event to obtain a target note event obtained by a musical instrument playing the target content; divide the notes contained in the target note event according to beats to obtain beat information to be analyzed; and obtain an analysis result for the musical instrument performance based on a matching result between the beat information to be analyzed and the beat template information of the target content.

[0016] The above-mentioned performance analysis method, electronic device, computer-readable storage medium and computer program product can identify the audio spectrum information obtained from the target content of the musical instrument performance to obtain a variety of note feature information, and then obtain the initial note event based on the multiple note feature information, and then filter the initial note event according to the note duration to obtain the target note event, thereby improving the reliability and accuracy of the obtained note event, and then divide the notes contained in the target note event according to the beat to obtain the beat information to be analyzed to ensure the reliability and accuracy of the beat information to be analyzed, and finally, according to the matching result of the beat information to be analyzed and the beat template information of the target content, the analysis result for the musical instrument performance can be obtained, thereby improving the reliability and accuracy of the performance analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 A diagram showing an application environment of a performance analysis method in one embodiment;

[0019] FIG2 (a) is a schematic diagram of a human voice spectrum;

[0020] FIG2( b ) is a schematic diagram of a piano performance spectrum;

[0021] Figure 3 A schematic diagram of a flow chart of a performance analysis method in one embodiment;

[0022] Figure 4 is a schematic diagram of multiple note feature information in one embodiment;

[0023] Figure 5 A schematic flow chart of the steps of obtaining an initial note event in one embodiment;

[0024] Figure 6 A schematic diagram of a sound head, a frame interval, and a sound tail in one embodiment;

[0025] Figure 7 is a schematic diagram of an initial note event in one embodiment;

[0026] Figure 8 is a schematic diagram of the start, sustain, decay, and release of a note corresponding to an embodiment;

[0027] Fig. 9 A schematic diagram of a note mapping result in one embodiment;

[0028] Fig.10 is a schematic diagram of beat information in one embodiment;

[0029] Fig.11 A flowchart of steps for obtaining analysis results for musical instrument performance in one embodiment;

[0030] Fig.12 is a schematic diagram of a matching result in one embodiment;

[0031] Fig.13 A schematic diagram of a performance analysis result in one embodiment;

[0032] Fig.14 A schematic diagram of a flow chart of a performance analysis method in another embodiment;

[0033] Fig.15 FIG. 4 is a diagram showing the internal structure of an electronic device in one embodiment. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0035] The performance analysis method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the application environment may include a terminal and a server, and the terminal communicates with the server through a network. The data storage system can store data that the server needs to process. The data storage system can be integrated on the server, or it can be placed on the cloud or other network servers. The performance analysis method provided in the embodiment of the present application can be executed by a terminal, or it can be executed by a server, or it can be executed by a terminal and a server in cooperation. Among them, the terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices, and the Internet of Things devices can be smart speakers, smart TVs, smart car devices, projection devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server can be an independent physical server, or it can be a server cluster or distributed system composed of multiple physical servers, or it can be a cloud server that provides cloud computing services.

[0036] In the current technology, artificial intelligence technology can be used to analyze songs sung by human voice, but this method is difficult to apply to the analysis of performances of related content using musical instruments. Taking the piano as an example, Figure 2 (a) and Figure 2 (b) show the human voice spectrum and the piano performance spectrum respectively. First, human voice singing is a single part, while piano performance is multiple parts. The fundamental frequency extraction algorithm for analyzing songs sung by human voice cannot handle piano performance. Secondly, the human voice sequence is one-dimensional, but the piano score is two-dimensional, that is, there will be multiple notes in the same beat. Thirdly, the human voice performance fluctuates greatly, but the piano notes are more stable. Based on this, based on the current related methods for analyzing songs sung by human voice, the performance analysis of related content using musical instruments will have technical problems such as inaccurate performance analysis.

[0037] The embodiment of the present application provides a performance analysis method that can be adapted to performances of related content using musical instruments, and can improve the reliability and accuracy of the performance analysis.

[0038] The performance analysis method of the present application is described below in conjunction with various embodiments and corresponding drawings.

[0039] In an exemplary embodiment, Figure 3 As shown, this method can be implemented as Figure 1 The terminal shown in the figure is executed, and the method may include the following steps:

[0040] Step S301: Acquire multiple note feature information obtained by recognizing the audio spectrum information of the target content.

[0041] The audio spectrum information is the audio spectrum information obtained by playing the target content with a musical instrument. The target content may be a music score or other playable content. The user may play the music score with a musical instrument such as a piano, and the terminal may collect the audio obtained by the user playing the music score with a musical instrument such as a piano, and convert the audio into audio spectrum information.

[0042] Among them, a variety of note feature information obtained by the audio transcription model from recognizing the audio spectrum information of the target content can be obtained. The audio transcription model can be an artificial intelligence model for audio transcription tasks. The current audio transcription model can be used to take audio spectrum information (such as Mel spectrum) as input and output a variety of note feature information.

[0043] Among them, the note feature information is information representing the note feature, and can be note feature information that can locate the position of the note. The note feature information can include multiple types. Based on the current audio transcription model, the audio spectrum information (such as Mel spectrum) can be used as input to output note feature information of types such as onset, offset, frame, and velocity. Taking the piano as an example, based on the current audio transcription model, the audio spectrum information (such as Mel spectrum) can be used as input to output note feature information of onset, offset, frame, and velocity. Each type of note feature information can include the probability of sounding of each key on the corresponding 88-key piano at the corresponding time point on the time axis. Among them, the note feature information of the type of onset can be recorded as onset feature information, the note feature information of the type of offset can be recorded as offset feature information, and the note feature information of the type of frame can be recorded as frame feature information.

[0044] As an example, Figure 4 A schematic diagram showing the general characteristics of these types of note features (the specific information contained in each diagram, such as the probability of sounding, is not the focus of this diagram). Figure 4 In the figure, from left to right, each column corresponds to four types of note features: onset, offset, frame, and velocity. The horizontal axis represents the time axis, and the vertical axis is the probability of each key on the 88-key piano being sounded at the timestamp corresponding to the corresponding time point.

[0045] In some embodiments of the present application, the terminal may obtain three types of note feature information, namely, onset, offset, and frame, obtained by an audio transcription model recognizing audio spectrum information.

[0046] Step S302: obtaining an initial note event of a target content played by a musical instrument according to the multiple note feature information, wherein the initial note event includes multiple notes.

[0047] In this step, one or more notes played when the target content is played at the corresponding time (such as one or more notes played at a certain time point) can be found according to the multiple note feature information in the order of the playing time to obtain the initial note event obtained by the instrument playing the target content. The initial note event can include each note played arranged according to the playing time obtained by this step.

[0048] Step S303: Filter the initial note event according to the note duration of the note contained in the initial note event to obtain a target note event obtained by performing the target content of the musical instrument.

[0049] In this step, the terminal can filter the initial note events according to the note duration of the notes contained in the initial note events, and can filter out the notes in the initial note events, such as notes with too short duration. The filtered initial note events can be used as target note events obtained as the target content of the instrument performance.

[0050] Step S304: Divide the notes included in the target note event according to beats to obtain beat information to be analyzed.

[0051] In this step, the terminal can divide the performance time into each beat according to a certain duration, that is, take a certain duration as one beat, and divide the notes contained in the target note event into a beat sequence with a certain number of beats according to the beats. Each beat can contain one or more notes, or may not contain any notes. The beat sequence obtained by division can be used as the beat information to be analyzed.

[0052] Step S305 , obtaining an analysis result for the musical instrument performance according to the matching result between the beat information to be analyzed and the beat template information of the target content.

[0053] The target content's beat template information can be used as an evaluation and analysis standard for the above-mentioned beat information to be analyzed. The target content's beat template information can be obtained from the target content's Musical Instrument Digital Interface (MIDI) file. For example, the note event template can be obtained based on the target content's Musical Instrument Digital Interface file, and the performance time can be divided into each beat according to the above-mentioned certain duration to obtain the target content's beat template information.

[0054] Based on this, the terminal can match the notes of each beat in the beat information to be analyzed with the notes of each beat in the beat template information of the target content (the matching method is not limited in this embodiment), and the matching result obtained can be used as the analysis result for the performance. The analysis result of the performance may include a score for the performance, the notes played correctly and their number, the notes missed and their number, the notes played incorrectly and their number, etc.

[0055] The performance analysis method of the present embodiment can identify the audio spectrum information obtained from the target content of the musical instrument performance to obtain a variety of note feature information, then obtain an initial note event based on the multiple note feature information, and then filter the initial note event based on the note duration to obtain a target note event, thereby improving the reliability and accuracy of the obtained note event, and then divide the notes contained in the target note event according to the beat to obtain the beat information to be analyzed to ensure the reliability and accuracy of the beat information to be analyzed, and finally, according to the matching result of the beat information to be analyzed and the beat template information of the target content, an analysis result for the musical instrument performance can be obtained, thereby improving the reliability and accuracy of the performance analysis.

[0056] In an exemplary embodiment, the multiple note feature information may include note head feature information, note tail feature information and inter-frame feature information; Figure 5 As shown, the initial note event of obtaining the target content of the musical instrument performance according to the multiple note feature information in step S302 may include:

[0057] Step S501, for each available note of the musical instrument, determine the note head position of the available note in the note head feature information, obtain the start timestamp of the available note according to the note head position, and obtain the end timestamp of the available note according to the start timestamp, the note tail feature information and the inter-frame feature information.

[0058] Among them, the available notes are notes of the musical instrument that can be used to play the target content. The available notes can be determined according to the musical instrument used. Taking the piano as an example, the available notes can include notes corresponding to 88 piano keys.

[0059] In this step, the note head position of the available note can be determined in the note head feature information according to the order of, for example, 88 piano keys (from top to bottom, in ordinate order) for each available note of the instrument, from the beginning to the end (in ordinate order) of the note head feature information. The note head position can be calibrated with a timestamp, and the note head position can be used to indicate at which point in the performance time the user played the note. The timestamp corresponding to the note head position can be used as the starting timestamp of the available note. Figure 6 After obtaining the starting timestamp of the available note, since the note head feature information, the note tail feature information and the inter-frame feature information correspond to the same time axis, the starting timestamp of the available note can be located in the note tail feature information and the inter-frame feature information respectively, and then the ending timestamp of the available note can be determined in the note tail feature information and the inter-frame feature information respectively according to the probability information of the note tail and the probability information between frames at the subsequent time of the starting timestamp.

[0060] Step S502, obtaining an initial note event according to the start timestamp and end timestamp of each available note.

[0061] In this step, each available note can be recorded on the time axis according to the start timestamp and the end timestamp of each available note. For example, the available note can be marked as a short line segment on the time axis according to the start timestamp and the end timestamp. Figure 7 After completing the corresponding identification of each available note on the timeline, the initial note event can be obtained.

[0062] The solution of this embodiment can quickly and effectively search for each available note of the instrument used to play the target content with the help of the head feature information, tail feature information and inter-frame feature information to obtain the initial note event.

[0063] Further, in one embodiment, determining the note head position of the available note in the note head feature information in step S501 may include:

[0064] The note head peak point of the available note is obtained from the note head feature information; the note head peak point is the time point in the note head feature information that satisfies the note head probability threshold condition and the monotonicity condition; the note head position of the available note is obtained according to the note head peak point.

[0065] This embodiment provides a method for accurately determining the sound head position of an available note. In this embodiment, a sound head probability threshold can be set. For the available notes, each sound head peak point that exceeds the sound head probability threshold and satisfies the monotonicity condition can be found in the sound head feature information in the order of the time axis, wherein the division standard of the sound head peak point can be a point with a large sound head probability at four adjacent points (time points) on the left and right of the time axis, for example, and satisfies the monotonicity condition of monotonically increasing on the left side of the point and monotonically decreasing on the right side of the point. Then the sound head peak point can be used as the sound head position of the available note.

[0066] Further, in one embodiment, obtaining the end timestamp of the available note according to the start timestamp, the end feature information and the inter-frame feature information in step S501 may include:

[0067] According to the inter-frame probability threshold condition, the inter-frame corresponding to the start timestamp is determined in the inter-frame feature information; and according to the end-of-sound probability threshold condition, the start timestamp and the end-of-sound position corresponding to the inter-frame corresponding to the start timestamp are determined in the end-of-sound feature information; if the duration represented by the timestamp corresponding to the end-of-sound position and the start timestamp meets the duration threshold condition, the timestamp corresponding to the end-of-sound position is determined as the end timestamp of the available note.

[0068] The solution of this embodiment can first determine the frame interval corresponding to the start timestamp in the inter-frame feature information according to the start timestamp, and then determine the corresponding sound tail position in the sound tail feature information according to the start timestamp and the frame interval corresponding to the start timestamp, and determine the end timestamp based on the timestamp corresponding to the sound tail position.

[0069] In this embodiment, an inter-frame probability threshold condition may be set, such as being lower than a certain inter-frame probability threshold. According to the inter-frame probability threshold condition, an inter-frame lower than a certain inter-frame probability threshold may be found based on the probability information of the inter-frame at the start timestamp at the subsequent time as the inter-frame corresponding to the start timestamp. Then, a sound tail probability threshold condition may be set, such as being higher than a certain sound tail probability threshold. According to the sound tail probability threshold condition, the position of the sound tail higher than a certain sound tail probability threshold may be found based on the probability information of the sound tail of the inter-frame corresponding to the start timestamp and the start timestamp at the subsequent time as the sound tail position corresponding to the inter-frame. Next, a duration threshold condition may be set, such as exceeding a certain duration threshold, and the duration represented by the timestamp corresponding to the sound tail position and the start timestamp is obtained. If the duration exceeds the duration threshold, the timestamp corresponding to the sound tail position may be determined as the end timestamp. Based on this, the scheme of this embodiment may combine the connection between the sound head feature information, the sound tail feature information and the inter-frame feature information and their respective probability information characteristics to determine the end timestamp based on the start timestamp.

[0070] In an exemplary embodiment, step S303 of filtering the initial note event according to the note duration of the note contained in the initial note event to obtain the target note event obtained by the instrument playing the target content may include:

[0071] Obtain a note duration threshold; the note duration threshold can be determined based on the shortest note duration of the target content; obtain a target note event based on the note whose note duration contained in the initial note event is greater than or equal to the note duration threshold.

[0072] In this embodiment, the note duration threshold can be determined based on the shortest note duration of the target content. The shortest note duration refers to the shortest duration among the note durations of the target content. As an example, the note duration threshold can be set using the duration corresponding to a sixteenth note as the shortest note duration.

[0073] In this embodiment, the terminal can filter the notes with a note duration less than the note duration threshold in the initial note event to obtain the target note event. The solution of this embodiment can filter the notes with too short a note duration in the initial note event according to the shortest note duration of the target content, thereby improving the accuracy and reliability of the target note event.

[0074] In an exemplary embodiment, the step S304 of dividing the notes included in the target note event according to beats to obtain the beat information to be analyzed may include:

[0075] The target note event is divided into a beat sequence according to the beat; and according to the mapping relationship between each available note and the target interval, the note corresponding to each beat in the beat sequence is mapped to the target interval to obtain the beat information to be analyzed. The available notes are the notes of the instrument that can be used to play the target content.

[0076] In this embodiment, in the target note event obtained through filtering, the timestamp corresponding to the head of each note contained therein can be used as the time point of the note to identify the time point of the note, so as to improve the accuracy of subsequent beat information matching. Figure 8 The ADSR (Attack, Decay, Sustain, Release) envelope shown in the figure has the strongest energy at the beginning of the note, and the entire energy change process increases monotonically. Finding the beginning of the note is equivalent to finding the peak point of energy to a certain extent. It is difficult to determine the attenuation of the end of the note, because, for example, the timbre of each piano is different, and the ADSR of each note is inconsistent. Among them, the energy at the position of the beginning of the note is basically a process of rapid construction from nothing to a clear boundary; the end of the note is to find the end point of energy attenuation, with a fuzzy boundary and more easily affected by background noise. The beginning of the note is the beginning of the playing, and is less affected by the echo; the end of the note is basically affected by the echo of the previous performance of the note. Considering the different energy attenuation curves of each frequency band, it is more likely to have an octave deviation. The characteristics between frames are similar to the end of the note.

[0077] In this embodiment, the target note event can be divided into a beat sequence according to a certain time interval based on the time point of the note. The beat of the beat sequence may contain one or more notes, or may not contain any notes. For example, the target note event can be divided into a beat sequence with a time interval of 10 milliseconds as one beat. The note corresponding to each beat in the beat sequence (which may contain one or more, or not) is also mapped to the target interval according to the mapping relationship to obtain the beat information to be analyzed. Among them, the mapping relationship is the mapping relationship between each available note and the target interval. The target interval can be an octave. The available notes are the notes of the musical instrument that can be used to play the target content. For example, the available notes corresponding to the 88 keys of the piano can be mapped to within an octave to obtain the beat information to be analyzed, so as to reduce the complexity of the performance analysis. In implementation, the available notes corresponding to the 88 keys can be mapped to within an octave by taking the modulo 12, such as Fig. 9 A schematic diagram showing a note mapping result.

[0078] In an exemplary embodiment, before obtaining the analysis result for the instrumental performance according to the matching result of the beat information to be analyzed and the beat template information of the target content in step S305, the method further includes:

[0079] The instrument digital interface file of the target content is obtained; and the beat template information is obtained according to the mapping relationship between the instrument digital interface file and each available note and the target interval, wherein the available notes are the notes of the instrument that can be used to play the target content.

[0080] In this embodiment, the beat template information of the target content can be obtained in the same manner as in the above embodiment. The terminal can obtain a MIDI file of the target content, and can obtain the beat template information of the target content based on the MIDI file and the mapping relationship between each available note and the target interval and the time interval.

[0081] Thus, the beat template information and the beat information to be analyzed can be binarized, and the binarized beat template information and the beat information to be analyzed can be used for subsequent matching processing. Fig.10 The figure in the figure is a schematic diagram showing binarized beat information, which can represent beat template information or beat information to be analyzed. The horizontal axis represents the time axis, and the vertical axis can represent notes within an octave.

[0082] The solution of this embodiment can convert the target note event into the beat information to be analyzed and obtain the beat template information of the target content according to the instrument digital interface file through a certain time interval and the mapping relationship between the available notes and the target interval, thereby reducing the complexity of the performance analysis.

[0083] Further, such as Fig.11 As shown, in one embodiment, the analysis result for the instrumental performance is obtained according to the matching result of the beat information to be analyzed and the beat template information of the target content in step S305, which may include:

[0084] Step S1101, performing first note matching on each beat in the beat information to be analyzed or a combination of adjacent beats corresponding to each beat and each beat in the beat template information to obtain a first matching result of the beat information to be analyzed.

[0085] Among them, the beat information to be analyzed can be matched with the beat template information in a dynamic programming manner to obtain the first matching result. In this step, considering the problem of recognition accuracy, some notes that should be in one beat may be split into two beats, so this step matches each beat in the beat information to be analyzed or the adjacent beat combination corresponding to each beat with each beat in the beat template information (this process is recorded as the first note matching), wherein the adjacent beat combination corresponding to each beat can be a beat combination formed by merging a beat with another beat adjacent to it before or after it, that is, it is allowed to match a beat in the beat information to be analyzed, and it is also allowed to merge two adjacent beats in the beat information to be analyzed to form a beat combination for matching. Among them, since the adjacent two beats as a beat combination have been matched, any beat in the beat combination is no longer split into one beat for matching again, so as to avoid repeated matching of a beat in the beat information to be analyzed, which affects the matching accuracy.

[0086] As a specific implementation method, in dynamic programming, a state transition equation can be defined:

[0087] A three-dimensional array dp is used, where dp[i][j][k] represents the minimum number of editing operations required to convert the first i elements of the user's beat sequence (beat information to be analyzed) into the first j elements of the beat template information, where k represents whether a combination operation is used (i.e., whether an adjacent beat combination is formed, k=0 indicates that the combination operation is not used, and k=1 indicates that the combination operation is used).

[0088] The state transfer equation can be expressed as:

[0089] 1. The jth beat of the beat template information is missed (missed): dp[i][j][0] = dp[i][j-1][0] +1;

[0090] 2. The user plays the i-th beat incorrectly (plays multiple times / plays incorrectly / plays multiple times): dp[i][j][0] = dp[i-1][j][0] + 1;

[0091] 3. The i-th beat of the user does not match the j-th beat of the beat template information (wrong play + missed play / multiple play / wrong performance / multiple performance + missed performance): dp[i][j][0] = dp[i-1][j-1][0] + (A[i-1] == B[j-1] ? 0 :1); where A represents an element in the user's beat sequence and B represents an element in the beat template information;

[0092] 4. Combined operation:

[0093] If i>1, you can consider combining A[i-2] and A[i-1] for matching: dp[i][j][1] = dp[i-2][j-1][0] + (combine(A[i-2],A[i-1]) == B[j-1] ? 0 : 1);

[0094] If i>2, you can consider combining A[i-3] and A[i-2] together for matching: dp[i][j][1] = dp[i-3][j-1][0] + (combine(A[i-3], A[i-2]) == B[j-1] ? 0 : 1).

[0095] The final dp[i][j][k] is the minimum value among the above operations.

[0096] The initial conditions can be:

[0097] -dp[0][j][0] = j: The user did not play anything, and the j-th beat template information is equal to missing j beats;

[0098] -dp[i][0][0] = i: The beat template information is empty, and the user plays i beats, which means that the user played i beats incorrectly;

[0099] -dp[0][j][1] = INF: The user does not play anything. In this case, the first j beats of the beat template information are not allowed to be combined.

[0100] -dp[i][0][1] = INF: The beat template information is empty, and the first i beats played by the user are not allowed to be combined.

[0101] Thus, through dynamic programming, each beat in the beat information to be analyzed or the combination of adjacent beats corresponding to each beat can be matched with each beat in the beat template information to obtain a first matching result of the beat information to be analyzed. The first matching result may include the matched beats and their number in the beat information to be analyzed, the matched beats and their number in the combination, the missed beats and their number, the wrongly played beats and their number, such as Fig.12 The matching result shown in can be used as the first matching result, wherein [o] represents a match, [8] represents a combination match, [?] represents multiple or wrong clicks by the user, and [x] represents missed clicks by the user.

[0102] Step S1102: if the first matching result includes a missed beat and the beat region where the missed beat is located does not include an erroneously played beat, then an analysis result is obtained according to the first matching result.

[0103] Among them, considering the problem of matching errors, the user's playing rhythm cannot be very accurate. In order to further improve the reliability and accuracy of matching, attention can be paid to the beat area that the user missed playing.

[0104] Based on this, the terminal can determine whether the first matching result contains a missed beat. If the first matching result contains a missed beat, the beat area where the missed beat is located can be determined, such as a beat area composed of one or more adjacent beats. Determine whether the beat area contains a wrongly played beat. That is, whether there is a beat / beat interval played by the user in the missed beat area.

[0105] In this step, if there is no beat played incorrectly by the user in the beat area that the user missed, that is, the beat area where the missed beat is located does not contain the incorrectly played beat, it can be determined that the user missed playing here, and the first matching result can be determined as the analysis result of the instrument performance.

[0106] Step S1103: If the first matching result includes a missed beat and the beat region where the missed beat is located includes an incorrectly played beat, a second note matching is performed on each beat in the beat information to be analyzed or a combination of adjacent beats corresponding to each beat with each beat in the beat template information to obtain a second matching result of the beat information to be analyzed, wherein the matching condition of the second note matching is different from the matching condition of the first note matching.

[0107] In this step, if there is a wrong beat in the beat area missed by the user, then the beat information to be analyzed needs to be note-matched with the beat template information again to obtain a second matching result of the beat information to be analyzed, wherein this matching is recorded as a second note match, and the matching condition of the second note match is different from the matching condition of the first note match. Compared with the first note match, the matching condition of the second note match can adopt a looser matching condition, such as matching one or more notes of several notes contained in a beat in the beat information to be analyzed with a note in the corresponding beat in the beat template information to be regarded as a hit, and so on.

[0108] In this step, if the second note matching still fails, such as it is still a missed note, it can be marked as a missed note. If the second note matching can be matched, it can be judged based on the number of wrongly played beats. If the wrongly played beat is a single beat, the beat can be left unmarked. If the wrongly played beat is multiple beats, the relaxed matching conditions of the second note matching can be used to mark the beat, such as [a beat is hit, but there is a wrong note / missed note] and so on.

[0109] Step S1104, obtaining an analysis result according to the second matching result.

[0110] In this step, the second matching result obtained after the second note matching can be used as the analysis result.

[0111] like Fig.13 As shown, taking the score of a performance as an example, the terminal can mark it in the score according to the analysis result as evaluation information of the user's performance, and assist the user in practicing instruments such as piano independently. Among them, the analysis results such as wrong playing, missed playing, and correct playing can be marked in the score in different ways, and there is no specific limitation on how to mark these results in the score.

[0112] The solution of this embodiment can accurately match the beat information to be analyzed with the beat template information of the target content, avoid problems such as user playing errors or rhythm deviations, and further improve the accuracy and reliability of performance analysis.

[0113] In one embodiment, in combination Fig.14 The overall framework and application scenarios of the performance analysis method of the present application are explained.

[0114] The performance analysis method of this embodiment can be applied to the evaluation and analysis of the performance behavior of piano beginners, for example, to make a reasonable evaluation of their performance, and can also assist beginners in practicing the piano independently by combining the annotation information of whether each beat of the music score is played correctly.

[0115] In this embodiment, the audio spectrum information obtained from the user's piano performance score can be input into the audio transcription model, and the audio transcription model can identify the audio spectrum information to obtain the sound start feature information, inter-frame feature information and frame end feature information, and obtain the target note event based on the sound start feature information, inter-frame feature information and frame end feature information. The target note event is then converted into beat information to be analyzed, and the instrument digitization interface file of the score is also obtained. The beat template information is obtained based on the instrument digitization interface file, and then the beat information to be analyzed and the beat template information can be dynamically planned and matched to obtain the performance analysis result of the user's performance score.

[0116] The performance analysis method of the present application can use an audio transcription model to identify the note head feature information, inter-frame feature information and frame tail feature information of the user's performance process, and based on this, it can combine the note duration screening, filtering and synthesis to obtain the target note event, thereby improving the accuracy and reliability of the performance analysis, and then divide the beats and map the target intervals according to the target note event to obtain the beat information to be analyzed, thereby reducing the complexity of the performance analysis, and finally dynamic planning matching can be used to obtain the analysis results of the user's piano performance, which may include information such as the correct beat played by the user, the beats played multiple times, the beats missed, and the evaluation score of the user's piano performance, etc.

[0117] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0118] In an exemplary embodiment, an electronic device is provided. The electronic device may be a terminal, and its internal structure diagram may be as shown in FIG. Fig.15 As shown. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and an external device. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a performance analysis method is implemented. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the electronic device casing, or an external keyboard, touchpad or mouse.

[0119] Those skilled in the art will understand that Fig.15 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0120] In one embodiment, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0122] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0123] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0124] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0125] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0126] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A performance analysis method, characterized in that: The method comprises: Acquire multiple note feature information obtained by identifying audio spectrum information of target content; the audio spectrum information is audio spectrum information obtained by playing the target content with a musical instrument; According to the multiple note feature information, an initial note event of a musical instrument playing the target content is obtained; the initial note event includes multiple notes; Filtering the initial note event according to the note duration of the note contained in the initial note event to obtain a target note event obtained by a musical instrument playing the target content; Dividing the notes contained in the target note event according to beats to obtain beat information to be analyzed; According to the matching result of the beat information to be analyzed and the beat template information of the target content, an analysis result for the instrument performance is obtained.

2. The method according to claim 1, characterized in that The multiple note feature information includes note head feature information, note tail feature information and inter-frame feature information; The obtaining, according to the plurality of note feature information, an initial note event obtained by a musical instrument playing the target content comprises: For each available note of the musical instrument, determine the note head position of the available note in the note head feature information, obtain the start timestamp of the available note according to the note head position, and obtain the end timestamp of the available note according to the start timestamp, the note tail feature information and the inter-frame feature information; wherein the available note is a note of the musical instrument that can be used to play the target content; The initial note event is obtained according to the start timestamp and the end timestamp of each of the available notes.

3. The method according to claim 2, characterized in that The step of determining the note head position of the available note in the note head feature information comprises: Acquire the note head peak point of the available note from the note head feature information; the note head peak point is the time point in the note head feature information that satisfies the note head probability threshold condition and the monotonicity condition; The note head position of the available note is obtained according to the note head peak point.

4. The method according to claim 2, characterized in that: The step of obtaining the end timestamp of the available note according to the start timestamp, the end feature information of the note, and the inter-frame feature information includes: Determine the interframe corresponding to the start timestamp in the interframe feature information according to the interframe probability threshold condition; and determine the start timestamp and the sound tail position corresponding to the interframe corresponding to the start timestamp in the sound tail feature information according to the sound tail probability threshold condition; If the duration represented by the timestamp corresponding to the end position of the note and the start timestamp meets the duration threshold condition, the timestamp corresponding to the end position of the note is determined as the end timestamp of the available note.

5. The method according to claim 1, characterized in that The filtering of the initial note event according to the note duration of the note contained in the initial note event to obtain the target note event obtained by the musical instrument playing the target content includes: Obtaining a note duration threshold; the note duration threshold is determined according to the shortest note duration of the target content; The target note event is obtained according to the note contained in the initial note event whose note duration is greater than or equal to the note duration threshold.

6. The method according to claim 1, characterized in that The step of dividing the notes contained in the target note event according to beats to obtain beat information to be analyzed includes: Dividing the target note event into a beat sequence according to the beat; According to the mapping relationship between each available note and the target interval, the note corresponding to each beat in the beat sequence is mapped to the target interval to obtain the beat information to be analyzed; wherein the available notes are the notes of the instrument that can be used to play the target content.

7. The method according to claim 1, characterized in that Before obtaining the analysis result for the instrumental performance according to the matching result between the beat information to be analyzed and the beat template information of the target content, the method further includes: Acquire a musical instrument digitization interface file of the target content; The beat template information is obtained according to the musical instrument digitization interface file and the mapping relationship between each available note and the target interval; wherein the available notes are the notes of the musical instrument that can be used to play the target content.

8. The method according to claim 1, characterized in that The obtaining of the analysis result for the instrument performance according to the matching result of the beat information to be analyzed and the beat template information of the target content includes: Performing first note matching on each beat in the beat information to be analyzed or a combination of adjacent beats corresponding to each beat with each beat in the beat template information to obtain a first matching result of the beat information to be analyzed; If the first matching result includes a missed beat and the beat region where the missed beat is located does not include an erroneously played beat, obtaining the analysis result according to the first matching result; If the first matching result includes a missed beat and the beat region where the missed beat is located includes an erroneously played beat, a second note matching is performed on each beat in the beat information to be analyzed or a combination of adjacent beats corresponding to each beat with each beat in the beat template information to obtain a second matching result of the beat information to be analyzed; the matching condition of the second note matching is different from the matching condition of the first note matching; The analysis result is obtained according to the second matching result.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.