A guitar intelligent teaching method and system and a storage medium

By analyzing audio and video data of guitar playing, we can identify and correct hand movements and rhythm problems, solve the irregularities in posture and chord conversion techniques in self-taught guitar, and achieve targeted adjustments and improvements.

CN116863787BActive Publication Date: 2025-10-10SHENZHEN MOOER AUDIO CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310136961.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-20
Publication Date
2025-10-10
Estimated Expiration
2043-02-20

AI Technical Summary

Technical Problem

When learning guitar by yourself, it is difficult to discover and correct problems with posture, chord change techniques, and rhythm and beat, which leads to irregular playing.

Method used

By recording audio and video data of guitar playing, using voice algorithms and hand movement recognition algorithms, the playing process is analyzed from both audio and video dimensions to identify and correct hand movement and rhythm problems.

Benefits of technology

Provide targeted adjustment suggestions to help users form good guitar playing standards and improve self-study results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116863787B_ABST
    Figure CN116863787B_ABST
Patent Text Reader

Abstract

The application relates to the field of guitar teaching, and specifically discloses a guitar intelligent teaching method, a system and a storage medium, the method comprising the following steps: acquiring input guitar playing data; performing feature extraction through a preset voice algorithm to acquire audio feature information; comparing the audio feature information with preset standard audio features to acquire a difference time interval set; extracting image sequence frames from video data based on the difference time interval set; performing feature extraction on the image sequence frames through a preset hand motion recognition algorithm to acquire a hand motion feature information set, comparing the hand motion feature information set with preset standard hand motion feature information to acquire an abnormal motion information set, and finally outputting corresponding data analysis information according to the current playing. The application detects and recognizes the guitar playing sound and hand motion from two dimensions of audio and video through the guitar playing data input by the user, finds out the problems in playing, and helps the user to adjust and improve in a targeted manner.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of guitar teaching, in particular to a guitar intelligent teaching method, system and storage medium. BACKGROUND

[0002] Guitar as a more popular musical instrument, is loved by more and more people, and more and more people begin to learn by watching relevant teaching materials, but this self-learning way is easy to go astray without a good foundation, for example, eager for success, too much pursuit of complete playing of the song, ignoring the correct playing posture, chord transformation skills and rhythm rhythm, etc.

[0003] Although there are relevant basic teaching in the teaching video watched, it can also bring certain help, but it is not as good as face-to-face guidance and correction by a teacher, and it is difficult to find one's own problems in the process of self-learning, so as to adjust in time. SUMMARY

[0004] The purpose of the present application is to provide a guitar intelligent teaching method, system and storage medium, which analyzes from two dimensions of audio and playing gesture action by recording the playing process, and points out the problems in the playing process.

[0005] In a first aspect, the present application provides a guitar intelligent teaching method, which adopts the following technical solution:

[0006] Obtain the recorded guitar playing data, and pre-process the playing data, wherein the guitar playing data includes audio data and video data;

[0007] Extract features of the audio data by a preset voice algorithm to obtain audio feature information;

[0008] According to the audio feature information, compare with a preset standard audio feature to obtain a difference time interval set;

[0009] Based on the difference time interval set, extract image sequence frames from the video data;

[0010] Extract features of the image sequence frames by a preset hand action recognition algorithm to obtain a hand action feature information set;

[0011] According to the hand action feature information set, compare with a preset standard hand action feature information to obtain an abnormal action information set;

[0012] According to the abnormal action information set, output corresponding data analysis information for the current playing.

[0013] Through the above technical solution, the guitar playing data recorded by the user can be detected and identified from the audio and video dimensions respectively, and based on the problems with the guitar playing sound, the hand movements during the playing process can be identified in a targeted manner to find out the incorrect or irregular hand movements, helping users to make timely adjustments during the self-study process.

[0014] Optionally, after obtaining the input guitar playing data and preprocessing the playing data, the method further includes:

[0015] Get the duration of the preprocessed audio data;

[0016] Determine whether the duration of the audio data is consistent with the duration of the preset standard audio data,

[0017] If not, extract the detection time segment set from the audio data by presetting the beat time segment set;

[0018] Perform beat recognition on the detection time segment set using a preset beat recognition method to obtain abnormal beat time segments;

[0019] Output corresponding prompt information and store the abnormal beat timing segment in the preset data storage space;

[0020] If so, the audio data is subjected to feature extraction using a preset voice algorithm to obtain audio feature information.

[0021] Optionally, after extracting features from the audio data using a preset voice algorithm to obtain audio feature information, the method further includes:

[0022] Identify audio feature information using preset key features to obtain the key and timbre of the current guitar playing;

[0023] Determine whether the currently played key and timbre are the same as the preset standard audio key and timbre;

[0024] If they are not the same, the audio feature information is converted into the same according to the current playing key, timbre and the preset standard audio key and timbre through the preset style transfer algorithm.

[0025] Optionally, the audio feature information includes time sequence information, and the step of performing comparison based on the audio feature information and obtaining a set of difference time intervals by using preset standard audio features includes:

[0026] Segment the audio feature information according to the time sequence information to obtain a segmented audio feature set;

[0027] Traverse the segmented audio feature set and calculate the similarity using the preset standard audio features;

[0028] When the similarity does not reach the preset threshold, the corresponding segmented audio feature is recorded as an abnormal segmented audio feature;

[0029] After the traversal is completed, all abnormal segmented audio features are obtained, and based on the corresponding timing information, a set of difference time intervals is obtained.

[0030] Optionally, after extracting the image sequence frames from the video data based on the difference time interval set, the method further includes:

[0031] According to the image sequence frames, key frames are extracted by a preset method;

[0032] According to the key frames, the posture estimation is obtained through the preset human posture recognition algorithm;

[0033] According to the posture estimation, a comparison is performed through a preset standard posture to obtain comparison result information, and the comparison result information is stored in a preset data storage space.

[0034] Optionally, the hand motion feature information set is compared with preset standard hand motion feature information to obtain an abnormal motion information set, including:

[0035] Traverse the hand motion feature information set, compare each frame of hand motion information with the standard hand motion feature information, and obtain the comparison result;

[0036] If the comparison results are inconsistent, the hand movement information of this frame is recorded as abnormal movement information;

[0037] After the traversal is completed, an abnormal action information set is generated for all acquired abnormal action information;

[0038] The abnormal action information set and the corresponding comparison results are stored in a preset data storage space.

[0039] Optionally, the abnormal action information includes chord mis-pressing information, and the abnormal action information set includes:

[0040] Traverse the abnormal action information set, determine the corresponding image data frame according to the chord wrong pressing information, and obtain the chord sequence frame and chord conversion sequence frame through the preset hand action recognition method;

[0041] Extract hand movement feature information based on the chord sequence frames and compare it with the preset standard chord fingering movements to obtain chord fingering error information;

[0042] Extracting hand movement change information based on the chord transition sequence frames, and comparing the hand movement information with the preset standard chord transition movements to obtain chord transition error information;

[0043] After the traversal is completed, all chord fingering error information and chord conversion error information are obtained and stored in a preset data storage space.

[0044] Optionally, after outputting corresponding data analysis information for the current playing according to the abnormal action information set, the method further includes:

[0045] Obtaining a data analysis overview of the current guitar playing based on the abnormal information recorded in the preset data storage space, wherein the data analysis overview includes a score of the current playing and problems;

[0046] Generate corresponding guidance suggestions based on the data analysis overview.

[0047] In a second aspect, the present application provides a guitar intelligent teaching system, comprising:

[0048] A data acquisition module (101) is used to acquire input guitar playing data and pre-process the playing data, wherein the guitar playing data includes audio data and video data;

[0049] An audio feature extraction module (102) is used to extract features from audio data using a preset voice algorithm to obtain audio feature information;

[0050] An audio feature comparison module (103) is used to compare the audio feature information with a preset standard audio feature to obtain a difference time interval set, and extract image sequence frames from the video data based on the difference time interval set;

[0051] A hand feature extraction module (104) is used to extract features from the image sequence frames using a preset hand action recognition algorithm to obtain a hand action feature information set;

[0052] A hand feature comparison module (105) is used to compare the hand motion feature information set with preset standard hand motion feature information to obtain an abnormal motion information set;

[0053] The data analysis module (106) is used to output corresponding data analysis information for the current playing according to the abnormal action information set.

[0054] In a third aspect, the present application provides a computer-readable storage medium storing a computer program that can be loaded by a processor and execute the above-mentioned intelligent guitar teaching method.

[0055] In summary, this application detects and identifies guitar playing sounds and hand movements from both audio and video dimensions in user-entered guitar playing data, allowing it to identify problems and their causes. Furthermore, it provides corresponding identification of the user's guitar playing movements and hand-holding posture, helping the user to adjust their posture to a certain extent and develop good guitar playing movement standards. Furthermore, it provides corresponding analysis data and guidance suggestions for playing problems, further helping users make targeted adjustments and improvements to achieve better self-study results. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is a flow chart of an intelligent guitar teaching method provided by an embodiment of the present application;

[0057] Figure 2 This is a flow chart of detecting the rhythm and beat of guitar playing provided by an embodiment of the present application;

[0058] Figure 3 This is a flowchart of obtaining a set of difference time intervals based on audio feature information provided by an embodiment of the present application;

[0059] Figure 4 This is a flowchart of obtaining abnormal action information set provided by an embodiment of the present application;

[0060] Figure 5 Schematic diagram of an intelligent guitar teaching system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] The following is combined with Figure 1 -Attached Figure 5 , further details of this application are given.

[0062] This application provides a guitar intelligent teaching method, see Figure 1 , including the following steps:

[0063] S100: Acquire recorded guitar playing data and pre-process the playing data, wherein the guitar playing data includes audio data and video data.

[0064] The audio data represents the sound of the guitar, and the video data represents dynamic frame data that records the user's posture and hand movements when playing the guitar.

[0065] Preprocessing means performing corresponding processing on audio data and video data. For example, preprocessing audio data will convert the acquired audio data into a specific file format and process it into a computer-recognizable binary feature vector form for storage. It will also perform noise reduction and frame and windowing on the audio data. Framing and windowing means dividing the continuous signal into independent frequency domain stable parts using acquisition windows of different lengths for easy analysis.

[0066] Preprocessing the video data will convert the video data into a sequence of frame images, and perform denoising and equalization on the sequence of frame images so that they can serve as data input for subsequent correlation image algorithms.

[0067] Because this application aims to help users identify problems and make timely adjustments in the process of self-learning guitar, it focuses on the user's guitar playing, such as the control of beat, whether the playing posture is standard, whether the chords are played correctly, whether the plucking and sweeping movements are correct, etc., so it does not take into account the user's singing, that is, it will not analyze or evaluate the user's singing.

[0068] In an embodiment of the present application, a recorded video of a user playing the guitar is obtained, audio data and video data are extracted from the video, and the audio data and video data are preprocessed accordingly, and the processed data are stored.

[0069] After obtaining audio data and video data, relevant explicit features can also be obtained. The so-called explicit features are information that can be directly obtained from the surface, such as the duration of audio data, the resolution size of video data, and other data, which can be reflected intuitively; while implicit features are information that requires certain processing or operations to be extracted, such as feature information extracted through relevant network models.

[0070] Therefore, after obtaining the audio data, the duration of the audio data can be obtained, and the duration is related to the rhythm and beat of the playing. Therefore, before subsequent processing of the audio data, you can start with the beat, that is, make a corresponding judgment on whether the user's grasp of the rhythm of playing the guitar is accurate. When the rhythm and beat are correct, you can better locate the problem of playing the melody, and thus find the cause of the problem in a targeted manner.

[0071] Therefore, in the embodiment of the present application, after obtaining the input guitar playing data and pre-processing the playing data, the guitar playing data is played. Figure 2 , including the following steps:

[0072] S110: Obtain the duration of the pre-processed audio data.

[0073] S120: Determine whether the duration of the audio data is consistent with the duration of preset standard audio data.

[0074] Among them, the preset standard audio data represents the audio data extracted from the existing standard guitar playing video data. The so-called standard guitar playing video data can be understood as the reference data of the guitar playing that the user is learning. It can be already in the preset database, or it can be the user finding relevant materials and entering it. However, the standard guitar playing video entered by the user himself needs to pass the corresponding detection program and undergo corresponding preprocessing before it can be used as a reference comparison standard data.

[0075] It's worth noting that the reason the audio data duration is checked after pre-processing is because the pre-processing also removes silence at the beginning and end of the audio data. This removes the silence at the beginning and end of the recorded audio data, taking into account the uncertainty of the recorded audio data. The remaining non-silent sound is used as the final audio data. This silence removal is determined by a pre-set algorithm, such as the VAD algorithm.

[0076] In an embodiment of the present application, after preprocessing the audio data, the corresponding duration information can be obtained. By comparing it with the duration of preset standard audio data, it can be known whether there is a rhythm problem in the current user's guitar playing. The consistency of duration here means that the difference in duration between the two is within the preset threshold range.

[0077] S130: If the duration of the audio data is inconsistent with the duration of the preset standard audio data, extract a detection time segment set from the audio data according to the preset beat time segment set.

[0078] S140 , performing beat recognition on the detection time segment set using a preset beat recognition method to obtain abnormal beat time segments.

[0079] S150: Output corresponding prompt information and store the abnormal beat timing segment in a preset data storage space.

[0080] The preset beat timing segment set represents the timing segments corresponding to the beats extracted from the preset standard guitar playing video.

[0081] In an embodiment of the present application, if it is determined that the duration of the current audio data is inconsistent with the corresponding preset standard audio data duration, it indicates that there is a problem with the guitar playing rhythm, so further specific analysis will be performed to find out the specific reasons that affect the rhythm error.

[0082] Since there are many variations in the rhythm of guitar playing, such as 2 / 4, 3 / 4, and 4 / 4, 2 / 4 means that each measure has two beats and a dynamic pattern of "strong and weak" with a quarter note as one beat; 3 / 4 means that each measure has three beats and a dynamic pattern of "strong, weak, weak" with a quarter note as one beat; and 4 / 4 means that each measure has four beats and a dynamic pattern of "strong, weak, secondary, strong, weak" with a quarter note as one beat.

[0083] Therefore, the corresponding segmented audio data will be extracted from the current audio data according to the preset beat timing segments, which will be recorded as the detection timing segment set. The detection timing segment set and the preset beat timing segment set will be compared one by one through the preset beat recognition method. The timing segments with different comparison results will be recorded as abnormal beat timing segments, and the abnormal beat timing segments will be stored in the preset data storage space.

[0084] S160: If the duration of the audio data is consistent with the duration of the preset standard audio data, feature extraction is performed on the audio data using a preset voice algorithm to obtain audio feature information.

[0085] In the embodiment of the present application, if the duration of the current audio data is determined to be consistent with the duration of the corresponding preset standard audio data, this does not indicate a lack of rhythm. Instead, it can be said that the timing is consistent with the standard guitar playing audio data, which is equivalent to the speed of the entire playing process. Further detailed comparative analysis can be performed, that is, using the corresponding voice algorithm to obtain audio feature information, and then comparing the audio data based on the audio feature information to determine whether there are any problems with the melody, timbre, pitch, etc.

[0086] S200: Extract features from audio data using a preset voice algorithm to obtain audio feature information.

[0087] The pre-set speech algorithms primarily include Linear Prediction Cepstral Coefficients (LPCC), Mel-Frequency Cepstral Coefficients (MFCC), and Constant Q Transform (CQT). Since the speech signal corresponding to the audio data is equivalent to a time-domain graph, or speech waveform, the speech algorithm's purpose is to filter the time-domain speech signal to obtain a nonlinear frequency-domain signal that is more suitable for extracting instrumental frequencies. Each waveform frame is then converted into a multidimensional vector containing sound information, which serves as audio feature information. This audio feature information represents a sequence of speech features extracted from the speech waveform that varies over time.

[0088] In an embodiment of the present application, after the audio data is acquired and preprocessed accordingly, feature extraction is performed on the audio data using a preset voice algorithm to obtain audio feature information.

[0089] Because the guitar also involves the issue of guitar key, such as C major, G major, etc., the standard sound produced by each string in different keys will also be correspondingly different. In addition, due to the different soundboard materials and string materials of different guitars, the timbre produced will also have certain differences.

[0090] Therefore, in order to further determine whether there are problems with the melody, audio change rhythm, tone, etc. of the guitar playing, it is first necessary to adjust the tone and timbre of the guitar playing to be consistent with the comparison standard.

[0091] In the embodiment of the present application, after extracting features from the audio data using a preset voice algorithm and obtaining audio feature information, the following steps are also included:

[0092] S210: Identify the audio feature information using a preset key feature to obtain the key and timbre of the current guitar playing.

[0093] S220: Identify the audio feature information using a preset base color feature to obtain the timbre of the current guitar playing.

[0094] S230: Determine whether the key and timbre of the current playing sound are the same as the preset standard audio key and timbre.

[0095] S240: If they are not the same, the audio feature information is converted into the same according to the currently played key, timbre and the preset standard audio key, timbre by using a preset style transfer algorithm.

[0096] The preset key features are the frequency domain feature information extracted from different keystrokes using a preset voice algorithm after the guitar is tuned. The standard audio key is the key obtained from the preset standard audio data, which can be used to understand the guitar key used by the teacher playing the song in the teaching video. The preset timbre features are the frequency domain feature information extracted from the preset voice algorithm, which can be understood as the guitar timbre used by the teacher playing the song in the teaching video.

[0097] In an embodiment of the present application, after obtaining audio feature information of the current audio data through a preset voice algorithm, the frequency domain features of the tone and the frequency domain features of the timbre can be extracted based on the audio feature information. By matching the frequency domain features of the current tone with the preset key features, the key of the current guitar playing can be identified. Similarly, by matching the frequency domain features of the current timbre with the preset timbre features, the timbre of the current guitar can be identified.

[0098] By comparing the key of the current guitar playing with the standard audio key, it can be determined whether the guitar key used is the same. By comparing the timbre of the current guitar playing with the standard audio timbre, it can be determined whether the guitar timbre used is the same. If there are differences, the audio feature information is converted into the same through the preset style transfer algorithm, that is, the current playing key and timbre are converted into the standard playing key and timbre in style to achieve consistency and reduce interference for subsequent audio feature information detection.

[0099] In addition, corresponding prompt information can be output to inform the user. If the key is different, the user can use a capo to change the key, or keep it unchanged according to their own style. If the timbre is different, the user will know that there may be some differences between the playing process and the one heard in the instructional video. This difference may be due to the different guitar timbre.

[0100] Because the fundamental key of a guitar is subjective during playing, and varies depending on the playing style, the user's pitch, and the combination with other instruments, there's generally no fixed key. However, determining the fundamental key of a guitar playing is crucial for identifying the subsequent guitar melody, timbre, and various high and low notes. Therefore, it's necessary to eliminate interference and maintain a consistent key when performing comparative analysis of audio data features.

[0101] S300: Based on the audio feature information, a set of difference time intervals is obtained by comparing with preset standard audio features.

[0102] Among them, the audio feature information contains timing information. As mentioned above, the voice signal corresponding to the audio data is equivalent to a time domain diagram. The waveform has a certain period in a short time. Different pronunciations often correspond to different periodic changes. Correspondingly, the audio feature information obtained by feature extraction of the audio data will also have this characteristic, so the corresponding time segment, that is, timing information, can be obtained based on this periodic change.

[0103] In an embodiment of the present application, audio feature information can be obtained by extracting features from audio data through a preset voice algorithm. The audio feature information is equivalent to a frequency domain change graph based on time changes. By comparing it with preset standard audio features, distinctive feature points can be obtained. For any distinctive feature point, a corresponding time interval can be generated, which is recorded as a difference time interval. Finally, a difference time interval set is obtained from all the distinctive feature points.

[0104] In the embodiment of the present application, based on the audio feature information, a set of difference time intervals is obtained by comparing with the preset standard audio features. Figure 3 , specifically including the following steps:

[0105] S310: Segment the audio feature information according to the time sequence information to obtain a segmented audio feature set.

[0106] S320: traverse the segmented audio feature set and perform similarity calculation using preset standard audio features.

[0107] S330: When the similarity does not reach a preset threshold, the corresponding segmented audio feature is recorded as an abnormal segmented audio feature.

[0108] S340: After the traversal is completed, all abnormal segmented audio features are obtained, and a difference time interval set is obtained according to the corresponding timing information.

[0109] In an embodiment of the present application, the timing information is equivalent to a set formed by several continuous time segments. The audio features can be segmented through the timing information to obtain a segmented audio feature set. For the same time segment, it can be compared through preset standard audio features, that is, by calculating the similarity between each segmented audio feature and the corresponding standard audio feature.

[0110] Since the audio feature information is equivalent to a set of multidimensional feature vectors, the similarity calculation is essentially to calculate the similarity between the multidimensional feature vectors of the audio feature information and the standard audio feature information. The similarity calculation results can be used to determine whether there are differences.

[0111] A similarity calculation is performed between each audio feature information and the corresponding standard audio feature information. If the similarity reaches a preset threshold, the two are considered similar. From the perspective of guitar playing, this means that the current playing segment meets the standard playing standards at the musical theory level. Since the guitar tone and timbre have been standardized, it can be simply understood here that the melody of the music played meets the playing standards.

[0112] If the similarity does not reach the preset threshold, it is considered that there is a difference between the two, that is, there is a certain problem with the current playing segment that does not meet the playing standard. Therefore, the segmented audio feature is recorded as an abnormal segmented audio feature, and the time sequence corresponding to the segmented audio feature is recorded as the difference time interval.

[0113] By traversing all the segmented audio features, the problematic segmented audio features and the corresponding difference time intervals can be obtained, and a difference time interval set is formed from all the difference time intervals for further detection to find the specific cause of the problem in playing.

[0114] S400 : Extracting image sequence frames from video data based on a set of difference time intervals.

[0115] By comparing the audio features of the current guitar playing with the preset standard audio features, we obtain corresponding distinguishing feature points and generate corresponding differential time intervals for these distinguishing feature points. Distinguishing feature points may be caused by incorrect chord pressing, incorrect chord conversion, non-standard strumming, or a lack of coordination with the chord pressing. Therefore, we need to conduct a detailed analysis of the current guitar playing video data based on the differential time intervals to identify the cause of the playing problem.

[0116] Therefore, in this embodiment of the present application, image sequence frames are extracted from the video data based on a set of differential time intervals. Since the video images contain dynamic frame data of the guitar playing posture and hand movements, the video data can be used to identify whether the hand movements during the playing process are standard, thereby discovering specific guitar playing problems.

[0117] It is worth noting that the reason why we first use audio data as a starting point is to extract abnormal time segments from the audio data, and then obtain corresponding image frame data based on the abnormal time segments.

[0118] This is because, first of all, the video data itself is converted into image frame data, which is a relatively large amount of data. In addition, each frame of image data often contains duplication and redundancy. Therefore, the corresponding image frame data is selected as a representative for detection. In addition, the standard recognition of hand movements during guitar playing is equivalent to a process of starting from a point and then expanding to the whole. For example, when determining whether the hand movements of plucking and sweeping the strings during guitar playing are standard, if the hand movements are not standard in a certain image frame data, it can be roughly determined that there is a problem with the user's guitar playing hand movements, without the need to identify and judge all the image frames data one by one.

[0119] In addition to hand gesture recognition, the system also identifies whether the fingers are plucking the strings correctly and at the appropriate timing. This is directly related to the audio data generated by the guitar, so audio data is used as a trigger for targeted video data detection. This allows for analysis of the specific causes of audio anomalies, which may include incorrect melody, incorrect pitch, or noise. Furthermore, it can identify the source of the error, namely whether there is a problem with the hand gesture or finger contact with the string.

[0120] Furthermore, since guitar playing posture—here referring to the way one holds the guitar while playing—has a standard position, whether standing or sitting, a significant deviation from the standard can affect the playing quality. Therefore, we can also detect the corresponding playing posture from video data.

[0121] In an embodiment of the present application, after extracting image sequence frames from video data based on the difference time interval set, the following steps are included:

[0122] S410 , extracting key frames from the image sequence frames using a preset method.

[0123] S420: Obtain posture estimation based on the key frame using a preset human posture recognition algorithm.

[0124] S430 , performing comparison based on the posture estimation using a preset standard posture, obtaining comparison result information, and storing the comparison result information in a preset data storage space.

[0125] The key frame refers to an image frame mentioned as a representative from the video data, and a preset method may be to select an image frame from the image sequence frames at a set time interval.

[0126] Since the image frames in the video data contain the user's guitar-playing posture, only representative image frames need to be selected for recognition. Usually, the user's guitar-holding posture will not change much during the guitar-playing process. Therefore, the image sequence frames obtained above can be further selected to obtain image frames that can detect the guitar-playing posture.

[0127] The preset human pose recognition algorithm can estimate the human pose by locating key points, for example, using the 3D Pose Estimation algorithm.

[0128] In an embodiment of the present application, the current guitar playing posture estimation can be obtained based on the key frame image through a preset human posture recognition algorithm, and the comparison result is obtained by comparing it with the preset standard posture. The comparison result contains specific difference information of the key points. If the difference is large, targeted adjustments can also be made based on the key point difference information, which is more conducive to forming a good guitar playing posture.

[0129] S500: Extract features from the image sequence frames using a preset hand motion recognition algorithm to obtain hand motion feature information.

[0130] Among them, the preset hand movement recognition algorithm uses a hand key point detection algorithm, such as the OpenPose hand key point detection model. The hand key point detection aims to locate the joint points and fingertip joint points on the fingers in a specified image.

[0131] Since a more comprehensive inspection and analysis will be conducted on problems encountered during guitar playing, especially the hand movements, multiple dimensions will be involved, such as whether the finger pressing of the strings is standard, whether the finger pressing of the strings is correct, whether the plucking and sweeping of the strings is standard, and the coordination of the two hands. If there are chords, it will also involve the issue of whether the chord transitions are wrong.

[0132] Therefore, in this embodiment of the application, a preset hand motion recognition algorithm is used to extract features from the image sequence frame data. This extracts hand motion feature information from each frame, namely hand key point information. This hand key point information can be used to calculate the curvature of the fingers, the spacing between the fingers, and the degree of contact between the fingers and the string. These data indicators can be compared with data from a standard playing video to identify abnormal data information.

[0133] S600: Obtain an abnormal motion information set by comparing the hand motion feature information set with preset standard hand motion feature information.

[0134] Among them, the preset standard hand movement feature information is the hand key point information extracted from the standard guitar playing video through the same hand movement recognition algorithm.

[0135] Since the standard guitar playing video corresponds to the current playing video, the user's current learning track is based on the standard playing video. Therefore, the hand key point data can be used to compare and judge whether the hand movements during guitar playing are correct or standard.

[0136] In the embodiment of the present application, the abnormal action information set is obtained by comparing the preset standard hand action feature information. Figure 4 , including the following steps:

[0137] S610 , traverse the hand motion feature information set, compare each frame of hand motion information with the standard hand motion feature information, and obtain a comparison result.

[0138] S620: If the comparison results are inconsistent, the hand motion information of the frame is recorded as abnormal motion information.

[0139] S630: After the traversal is completed, generate an abnormal action information set for all acquired abnormal action information.

[0140] S640: Store the abnormal action information set and the corresponding comparison result in a preset data storage space.

[0141] In an embodiment of the present application, a set of hand movement feature information can be obtained by performing feature extraction on an image frame sequence, that is, corresponding hand feature information can be obtained for each frame of image.

[0142] By traversing the hand motion feature information set, each frame's hand motion feature information can be compared with preset standard hand motion feature information. If the comparison results are inconsistent, it indicates a corresponding anomaly, and the hand motion information corresponding to the current frame is recorded as abnormal motion information. The abnormal motion information here is divided into corresponding categories based on the comparison results, such as the left hand finger pressing the wrong string, the right hand plucking rhythm pattern error, the consistency of the two hands' movements not consistent with standard playing, and the left hand pressing the wrong chord, or the chord is pressed correctly but not in accordance with the standard, resulting in a deviation in the guitar sound.

[0143] All hand motion feature information in the hand motion feature information set is traversed and compared one by one. After the traversal is completed, all abnormal motion information can be obtained and stored in a preset data set, which is recorded as the abnormal motion information set. At the same time, the abnormal motion information set is stored in the preset data storage space.

[0144] Because abnormal actions not only detect whether the finger pressing and strumming movements are standard, and whether the finger pressing is correct, but also detect whether the chord is pressed correctly and whether the chord transitions meet the standards, abnormal action information also includes chord mispressing information.

[0145] Therefore, in the embodiment of the present application, after generating the abnormal action information set, the following steps are also included:

[0146] S650: traverse the abnormal action information set, determine the corresponding image data frame according to the chord wrong pressing information, and obtain the chord sequence frame and the chord conversion sequence frame through a preset hand action recognition method.

[0147] The reason for pressing the wrong chord may be that you are not familiar with the chord, pressing the wrong position, or there is a problem with the chord transition, the timing of the transition is wrong or the transition speed is too slow.

[0148] Therefore, in the embodiment of the present application, further detection will be performed on the chord mis-pressing information.

[0149] First, by traversing the abnormal action information set, all the chord mis-pressing information can be obtained. Based on the chord mis-pressing information, all the image data frames of the chord mis-pressing can be determined. Then, the continuous image sequence frames containing the chord finger pressing can be extracted. The extraction is mainly performed through the preset hand movement recognition method. When the chord needs to be converted, the corresponding hand movement will also make corresponding changes. Therefore, through the preset hand movement recognition method, the chord sequence frames corresponding to the chord mis-pressing information and the core conversion sequence frames can be obtained.

[0150] S660: Extract hand movement feature information based on the chord sequence frames, and compare it with a preset standard chord fingering movement to obtain chord fingering error information.

[0151] In an embodiment of the present application, for a chord sequence frame, corresponding hand movement feature information can be obtained through a preset hand movement recognition method. By comparing the hand movement feature information with a preset standard chord fingering movement, it can be determined whether the chord fingering is correct. If not, chord fingering error information is obtained. The chord fingering error information specifically indicates the name of the chord currently being fingered incorrectly.

[0152] S670: Extract hand movement change information based on the chord conversion sequence frames, and compare the hand movement information with a preset standard chord conversion movement to obtain chord conversion error information.

[0153] The chord sequence frames represent the chords before and after the chord conversion and the image frames during the conversion process.

[0154] Therefore, in the embodiment of the present application, the chord transition sequence frames can also be used to extract hand motion feature information and hand motion change information based on a preset hand motion recognition method. By comparing the hand motion change information with a preset standard chord transition motion, chord transition error information can be obtained. In addition, because the chord sequence frames also include the chords before and after the chord transition, the chord being transitioned can be further identified, which also helps users practice and improve in a targeted manner.

[0155] S680: After the traversal is completed, all chord fingering error information and chord conversion error information are obtained and stored in a preset data storage space.

[0156] In the embodiment of the present application, by identifying and comparing all chord sequence frames and chord conversion sequence frames one by one, all chord fingering error information and chord conversion error information can be obtained, and all relevant error information can be stored in a preset data storage space.

[0157] S700: Output corresponding data analysis information for the current playing according to the abnormal action information set.

[0158] In an embodiment of the present application, the abnormal action information set represents abnormal information of hand movements, and the current guitar playing can be scored by analyzing the abnormal action information. In addition, corresponding information is output for specific problems so that the user can make targeted adjustments based on the problems in the playing.

[0159] Furthermore, since the acquired abnormal action information set is stored in a preset data storage space, in addition to the abnormal action information set, the preset data storage space also stores abnormal beat timing information, chord fingering error information, and chord conversion error information. Therefore, further analysis of the various abnormal information in the preset data storage space can be performed to provide targeted guidance and suggestions.

[0160] Therefore, in the embodiment of the present application, after outputting corresponding data analysis information for the current playing according to the abnormal action information set, the following steps are also included:

[0161] S710: Obtain a data analysis overview of the current guitar playing according to the abnormal information recorded in the preset data storage space.

[0162] S720: Generate corresponding guidance suggestions based on the data analysis overview.

[0163] The data analysis overview includes the score of the current performance and where the problems lie.

[0164] In the application embodiment, the abnormal information recorded in the preset data storage space can be used to obtain the overview of the current guitar playing, such as beat problems, in which time periods they occur, chord F failing to follow the specifications, errors in the transition from chord F to chord G, etc.

[0165] Then, based on the data analysis overview of the current guitar playing, corresponding guidance suggestions can be generated. For example, for beat problems, practice plans can be generated based on specific beat notes; for chord mis-pressing problems, the finger-pressing specifications for the chord can be broken down and the corresponding movements can be demonstrated, so that users can clearly understand how to press the chords well; for chord conversion problems, corresponding chord conversion practice plans can also be generated, so that users can adjust their practice in a targeted manner.

[0166] The present application also provides a guitar intelligent teaching system. Figure 5 The system includes: a data acquisition module 101, an audio feature extraction module 102, an audio feature comparison module 103, a hand feature extraction module 104, a hand feature comparison module 105, and a data analysis module 106.

[0167] The data acquisition module 101 is used to acquire input guitar playing data and pre-process the playing data, wherein the guitar playing data includes audio data and video data.

[0168] The audio feature extraction module 102 is used to extract features from audio data using a preset voice algorithm to obtain audio feature information.

[0169] The audio feature comparison module 103 is configured to perform comparison based on the audio feature information using preset standard audio features to obtain a set of difference time intervals, and extract image sequence frames from the video data based on the set of difference time intervals.

[0170] The hand feature extraction module 104 is configured to extract features from the image sequence frames using a preset hand motion recognition algorithm to obtain a hand motion feature information set.

[0171] The hand feature comparison module 105 is configured to compare the hand motion feature information set with preset standard hand motion feature information to obtain an abnormal motion information set.

[0172] The data analysis module 106 is used to output corresponding data analysis information for the current playing according to the abnormal action information set.

[0173] In the embodiment of the present application, the data acquisition module 101 is specifically used to acquire the recorded guitar playing data and pre-process the audio data and video data in the playing data respectively.

[0174] The audio feature extraction module 102 is specifically configured to extract features from audio data using a preset speech algorithm to obtain audio feature information.

[0175] The audio feature comparison module 103 is specifically configured to perform comparison based on the audio feature information using preset standard audio features to obtain a set of difference time intervals, and extract image sequence frames from the video data using the set of difference time intervals.

[0176] The hand feature extraction module 104 is specifically configured to extract features from the image sequence frames using a preset hand motion recognition algorithm to obtain a hand motion feature information set.

[0177] The hand feature comparison module 105 is specifically used to compare the hand movement feature information set with the preset standard hand movement feature information, obtain the abnormal movement information set, further detect the chord-related problems in the abnormal movement information, and store the relevant abnormal problems in the corresponding data storage space.

[0178] The data analysis module 106 is specifically used to perform data analysis on the current playing according to the abnormal action information set and various abnormal information in the data storage space, and output corresponding problem information and guidance suggestions.

[0179] An embodiment of the present application also provides a computer-readable storage medium storing a computer program that can be loaded by a processor and execute any of the above-mentioned intelligent guitar teaching methods.

[0180] The embodiments of the present application are preferred embodiments of the present application, and are not intended to limit the protection scope of the present application, and thus: any equivalent changes made according to the principles of the present application should be covered within the protection scope of the present application.

Claims

1. An intelligent guitar teaching method, characterized in that: include: Acquire recorded guitar playing data and pre-process the playing data, wherein the guitar playing data includes audio data and video data; Extract features from audio data using a preset voice algorithm to obtain audio feature information; Based on the audio feature information, a set of difference time intervals is obtained by comparing with preset standard audio features; Extracting image sequence frames from video data based on a set of difference time intervals; Perform feature extraction on the image sequence frames using a preset hand movement recognition algorithm to obtain a hand movement feature information set; According to the hand movement feature information set, the hand movement feature information set is compared with the preset standard hand movement feature information to obtain the abnormal movement information set; According to the abnormal action information set, the corresponding data analysis information is output for the current playing; Based on the hand motion feature information set, the abnormal motion information set is obtained by comparing it with the preset standard hand motion feature information, including: Traverse the hand motion feature information set, compare each frame of hand motion information with the standard hand motion feature information, and obtain the comparison result; If the comparison results are inconsistent, the hand movement information of this frame is recorded as abnormal movement information; After the traversal is completed, an abnormal action information set is generated for all acquired abnormal action information; Storing the abnormal action information set and the corresponding comparison results in a preset data storage space; The abnormal action information includes chord mis-pressing information. After generating the abnormal action information set, the method includes: Traverse the abnormal action information set, determine the corresponding image data frame according to the chord wrong pressing information, and obtain the chord sequence frame and chord conversion sequence frame through the preset hand action recognition method; Extract hand movement feature information based on the chord sequence frames and compare it with the preset standard chord fingering movements to obtain chord fingering error information; Extracting hand movement change information based on the chord transition sequence frames, and comparing the hand movement information with a preset standard chord transition action to obtain chord transition error information; After the traversal is completed, all chord fingering error information and chord conversion error information are obtained and stored in a preset data storage space.

2. A guitar intelligent teaching method according to claim 1, characterized in that: The method of obtaining the input guitar playing data and preprocessing the playing data includes: Get the duration of the preprocessed audio data; Determine whether the duration of the audio data is consistent with the duration of the preset standard audio data, If not, extract the detection time segment set from the audio data by presetting the beat time segment set; Perform beat recognition on the detection time segment set using a preset beat recognition method to obtain abnormal beat time segments; Output corresponding prompt information and store the abnormal beat timing segment in the preset data storage space; If so, the audio data is subjected to feature extraction using a preset voice algorithm to obtain audio feature information.

3. The intelligent guitar teaching method according to claim 1, wherein: After extracting features from the audio data using a preset voice algorithm to obtain audio feature information, the method further includes: Identify audio feature information using preset key features to obtain the key and timbre of the current guitar playing; Determine whether the currently played key and timbre are the same as the preset standard audio key and timbre; If they are not the same, the audio feature information is converted into the same according to the current playing key, timbre and the preset standard audio key and timbre through the preset style transfer algorithm.

4. The intelligent guitar teaching method according to claim 1, wherein: The audio feature information includes time sequence information, and the step of comparing the audio feature information with a preset standard audio feature to obtain a set of difference time intervals includes: Segment the audio feature information according to the time sequence information to obtain a segmented audio feature set; Traverse the segmented audio feature set and calculate the similarity using the preset standard audio features; When the similarity does not reach the preset threshold, the corresponding segmented audio feature is recorded as an abnormal segmented audio feature; After the traversal is completed, all abnormal segmented audio features are obtained, and the difference time interval set is obtained based on the corresponding timing information.

5. The intelligent guitar teaching method according to claim 1, wherein: After extracting the image sequence frames from the video data based on the difference time interval set, the method includes: According to the image sequence frames, key frames are extracted by a preset method; According to the key frames, the posture estimation is obtained through the preset human posture recognition algorithm; According to the posture estimation, a comparison is performed through a preset standard posture to obtain comparison result information, and the comparison result information is stored in a preset data storage space.

6. The intelligent guitar teaching method according to claim 1, characterized in that: After outputting corresponding data analysis information for the current playing according to the abnormal action information set, the method includes: Obtaining a data analysis overview of the current guitar playing based on the abnormal information recorded in the preset data storage space, wherein the data analysis overview includes a score of the current playing and problems; Generate corresponding guidance suggestions based on the data analysis overview.

7. A guitar intelligent teaching system using the guitar intelligent teaching method according to any one of claims 1 to 6, characterized in that: include: A data acquisition module (101) is used to acquire input guitar playing data and pre-process the playing data, wherein the guitar playing data includes audio data and video data; An audio feature extraction module (102) is used to extract features from audio data using a preset voice algorithm to obtain audio feature information; An audio feature comparison module (103) is used to compare the audio feature information with a preset standard audio feature to obtain a difference time interval set, and extract image sequence frames from the video data based on the difference time interval set; A hand feature extraction module (104) is used to extract features from the image sequence frames using a preset hand action recognition algorithm to obtain a hand action feature information set; A hand feature comparison module (105) is used to compare the hand motion feature information set with preset standard hand motion feature information to obtain an abnormal motion information set; The data analysis module (106) is used to output corresponding data analysis information for the current playing according to the abnormal action information set.

8. A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing the intelligent guitar teaching method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Musical instrument auxiliary teaching method and system based on AR technology

    CN114783222A