An intelligent recognition method for piano fingering

Through deep learning technology, combined with significance target detection and manual key point detection, intelligent identification and real-time correction of piano player fingering is achieved, and the problem of difficult to judge multiple player fingering in real time and correct errors by self-study in the existing technology is solved, which improves the efficiency and accuracy of piano learning.

CN112488047BActive Publication Date: 2025-05-30SHANGHAI ULUCU ELECTRON TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011482561.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-16
Publication Date
2025-05-30
Estimated Expiration
2040-12-16

AI Technical Summary

Technical Problem

In piano playing learning, it is difficult for the prior art to achieve real-time fingering judgments and error reminders for multiple players, especially for self-study students, it is difficult to correct errors without guidance.

Method used

Using an intelligent recognition method based on deep learning, the piano keyboard and player's fingers are monitored in real time through the camera, and using the significance object detection algorithm and the manual key point detection algorithm are used to identify and match the playing fingering, giving error prompts in real time.

Benefits of technology

It realizes intelligent identification and real-time correction of piano player fingering, facilitates the process of self-study piano, improves the efficiency of teaching guidance, and enhances the system's anti-interference ability to the external environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112488047B_ABST
    Figure CN112488047B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent method for identifying piano fingering, which includes the following steps: Camera installation and debugging: Segment the piano keyboard using a saliency object detection algorithm based on deep learning; Piano keyboard calibration: Detect the hand playing the piano; Detect the key point coordinates of the playing fingers using a hand key point detection algorithm; Joint matching of the information of the playing fingers and the piano keyboard. The method based on the deep convolutional neural network of the present invention usually designs a neural network model to mine deeper and more abstract image features, without manual participation, is less affected by factors such as illumination and posture, and has a stronger adaptability to complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for intelligent recognition of piano fingering. Background Art

[0002] In the scene of playing a piano score, it is necessary to manually confirm whether the playing technique is correct. In actual operation, this requires the observer to watch the finger movements of the player and be familiar with the score throughout the process. This makes it impossible for an observer to simultaneously observe multiple players and give timely judgments on the playing fingering. Especially for those who study by themselves without an instructor, it is simply impossible to study alone.

[0003] Piano playing learning mainly focuses on learning to play with fingers according to the score. In piano playing teaching, the correctness of playing is measured by judging whether the score and the corresponding playing fingering match one by one. In piano practice, it is almost impossible for a teacher to personally check whether the playing fingering of multiple students playing the same piano score or different piano scores at the same time is correct. With the development and progress of computer vision technology and machine learning, it has become increasingly possible to automatically recognize events through monitoring, such as the recognition of body movements, gesture recognition, face recognition, etc. These intelligent recognition methods basically extract the features of objects first, and then perform detection, classification, and recognition based on the features. The methods for extracting the features of objects are mainly divided into the method of traditional manual design features and the method based on deep convolutional neural networks:

[0004] 1. Methods of traditional manual design features such as HOG, LBP, SIFT, etc. The traditional manual design features are relatively simple, without the need for learning and training, only simple calculation and statistics are required. However, the manually designed features are easily affected by external factors, and the actual application effect is not robust.

[0005] 2. The method based on deep convolutional neural networks usually designs a neural network model to mine deeper and more abstract image features, without manual participation, less affected by factors such as illumination and posture, and has stronger adaptability to complex scenes. Summary of the Invention

[0006] The object of the present invention is to propose a method for intelligent recognition of piano fingering, which is aimed at intelligent recognition of errors in piano playing techniques during the process of learning to play the piano score alone and simultaneously evaluating multiple students playing the piano score, and can timely remind of error information, facilitating the player to receive timely error reminders and evaluations even without an instructor beside.

[0007] The object of the present invention is to solve the problem that when a piano player plays a score, by intelligently recognizing the playing fingering, timely prompt of error information of the playing fingering of the player is given, enabling the player to correct errors in time during practice and achieving the purpose of learning to play the piano alone.

[0008] The specific technical solution of the present invention is as follows:

[0009] An intelligent recognition method for piano fingering includes the following steps:

[0010] Step 1: Camera installation and debugging: Install a camera whose viewing angle can fully cover the piano keyboard, with clear picture quality as much as possible, and the camera can be connected to the piano display screen to display the piano keyboard picture on the piano screen in real time.

[0011] Step 2: Piano keyboard segmentation: For piano keyboard segmentation, the saliency object detection algorithm SOD100K[1] based on deep learning is mainly used. The lightweight network proposed by this algorithm is mainly composed of a feature extractor and an inter-stage fusion part, which can process features of multiple scales at the same time. The feature extractor is stacked with the intra-scale multi-scale blocks proposed by SOD100K and is divided into 4 stages according to the resolution of the feature map. Each stage has 3, 4, 6, and 4 intra-scale multi-scale blocks respectively. The inter-stage fusion part composed of a flexible convolutional module (gOctConvs) proposed by SOD100K will process the features from each stage of the feature extractor to obtain a high-resolution output.

[0012] This algorithm uses a new type of dynamic weight decay scheme to reduce the redundancy of feature representation and can adjust the weight decay according to the specific features of certain channels. Specifically, during backpropagation, the decay term will change dynamically according to the features of certain channels. The weight update of dynamic weight decay can be expressed as:

[0013]

[0014] where λ d is the weight of dynamic weight decay, x i represents the feature calculated by w i , and S(x i ) is the measure of the feature, which can have multiple definitions according to the task. w i is the weight of the i-th layer, is the gradient to be updated. In this algorithm, the goal is to allocate weights according to the features between stable channels, and global average pooling is used as an index for specific channels. The formula can be expressed as:

[0015]

[0016] x i represents the feature map, and H and W represent the height and width of the feature map respectively.

[0017] Step 3: Piano keyboard calibration: Through the box coordinates of the segmented keyboard obtained in Step 2, it is now necessary to sort the keyboard from left to right and calibrate the keyboard boxes such as X1, X2,....

[0018] Step 4: Detection of the human hand playing the piano: Collect a part of the finger video of playing the piano through a camera and collect a part of human hand pictures on the Internet, and label the human hand box and the left and right classification marks. Use the FaceBoxes[2] detection algorithm to train the human hand detector model. This algorithm proposes a new anchor box density increase strategy, aiming to improve the recall rate of small-scale human faces. Set anchor boxes on different feature maps to detect target objects. However, for the case of crowded targets, the small anchors set at the bottom layer of the network are obviously very sparse. In order to densify those small anchors at the bottom layer, specifically, at the center of each receptive field, offset it. The anchor density can be expressed as:

[0019] A density = A scale / A interval

[0020] A scale represents the scale of the anchor, while A interval represents the interval of the anchor.

[0021] Step 5: Detection of the key points of the playing fingers: Use the OpenPose[3] human hand key point detection algorithm to detect the coordinates of the key points of the playing fingers, and mark the key points close to the fingertips as h1, h2, h3,.... The network structure of this algorithm contains 6 stages. The loss of each stage is the L2 norm between the predicted values of the limb position confidence map and the limb affinity vector field and the ground truth, which can be expressed as:

[0022]

[0023]

[0024] are the predicted value and the true value of the limb position confidence map respectively, are the predicted value and the true value of the limb affinity vector field respectively. W(p) is 0 or 1. When it is 0, it means that the annotation of a certain key point is missing, and the loss does not calculate this point.

[0025] The overall loss is the sum of the losses of each stage:

[0026]

[0027] Step 6: Joint matching of playing finger and piano keyboard information: When the player presses the piano key X1, the piano side will send out a signal f1 corresponding to the pressed piano key. The signal f1 corresponds to the X1 frame of the piano key. Detect whether there is a human hand on the piano keyboard through Step 4. If there is, detect the key points of the human hand through Step 5, and compare whether each key point close to the fingertip falls within the piano key frame calibrated by X1. If there is a calibrated fingertip key point within the frame, the finger used by the player can be determined through the fingertip key point, and then it can be identified whether the playing fingering is correct. If it is incorrect, an error prompt will be given.

[0028] The intelligent recognition of fingering during piano playing can timely remind the player of incorrect fingering when playing the music score, timely correct and remind during solo piano practice, and assist in multiple piano teaching instructions at the same time, improving the efficiency of teaching instructions. It is convenient for the player to receive timely error reminders and evaluations even without a tutor beside. Using deep learning methods, models for piano keyboard segmentation, detection of the playing human hand, and detection of playing finger key points are highly resistant to external environmental interference and have good model robustness. Brief Description of the Drawings

[0029] Figure 1 It is a schematic diagram of the intelligent fingering recognition method for the piano of the present invention. Detailed Embodiment

[0030] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0031] An intelligent fingering recognition method for a piano, the overall steps are as follows.

[0032] Step 1: Installation and debugging of the camera: Install a camera whose viewing angle can fully cover the piano keyboard, with clear image quality as much as possible, and the camera can be connected to the piano display screen to display the piano keyboard image on the piano screen in real time.

[0033] Step 2 Piano Keyboard Segmentation: For piano keyboard segmentation, the saliency object detection algorithm SOD100K [1] (Highly Efficient Salient Object Detection with 100K Parameters) based on deep learning is mainly used. The lightweight network proposed by this algorithm mainly consists of a feature extractor and a cross-stage fusion part, which can process features of multiple scales simultaneously. The feature extractor is stacked with the intra-scale multi-scale blocks proposed by SOD100K and is divided into 4 stages according to the resolution of the feature map. Each stage has 3, 4, 6, and 4 intra-scale multi-scale blocks respectively. The cross-stage fusion part composed of a flexible convolutional module (gOctConvs) proposed by SOD100K processes the features from each stage of the feature extractor to obtain a high-resolution output.

[0034] This algorithm uses a new type of dynamic weight decay scheme to reduce the redundancy of feature representation and can adjust the weight decay according to the specific features of certain channels. Specifically, during backpropagation, the decay term changes dynamically according to the features of certain channels. The weight update of dynamic weight decay can be expressed as:

[0035]

[0036] where λ d is the weight of dynamic weight decay, x i represents the feature calculated by w i , and S(x i ) is the metric of the feature, which can have multiple definitions according to the task. w i is the weight of the i-th layer, is the gradient to be updated. In this algorithm, the goal is to allocate weights according to the features between stable channels, and global average pooling is used as an index for specific channels. The formula can be expressed as:

[0037]

[0038] x i represents the feature map, and H and W represent the height and width of the feature map respectively.

[0039] Step 3 Piano Keyboard Calibration: Given the bounding box coordinates of the segmented keyboard obtained in Step 2, it is now necessary to sort the keyboard from left to right and calibrate the keyboard bounding boxes as X1, X2,...

[0040] Step 4. Human hand detection: Collect a part of the finger video of playing the piano through a camera and collect a part of human hand pictures on the Internet, and label the human hand box and the left and right classification marks. Use the FaceBoxes[2] (FaceBoxes: A CPU Real-time Face Detector with High Accuracy) detection algorithm to train the human hand detector model. This algorithm proposes a new anchor box density increasing strategy, aiming to improve the recall rate of small-scale faces. Anchor boxes are set on different featuremaps to detect target objects. However, for the case of crowded targets, the small anchors set at the bottom layer of the network are obviously very sparse. In order to densify those small anchors at the bottom layer, specifically, at the center of each receptive field, it is offset. The anchor density can be expressed as:

[0041] A density = A scale / A interval

[0042] A scale represents the scale of the anchor, while A interval represents the interval of the anchor.

[0043] Step 5. Finger key point detection: Use the OpenPose[3] (OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields) human hand key point detection algorithm to detect the coordinates of the key points of the playing fingers, and mark the key points close to the fingertips as h1, h2, h3,.... The network structure of this algorithm contains 6 stages. The loss of each stage is the L2 norm between the predicted values of the limb position confidence map and the limb affinity vector field and the ground truth, which can be expressed as:

[0044]

[0045]

[0046] are the predicted value and the true value of the limb position confidence map respectively, are the predicted value and the true value of the limb affinity vector field respectively. W(p) is 0 or 1. When it is 0, it means that the annotation of a certain key point is missing, and the loss does not calculate this point.

[0047] The overall loss is the sum of the losses of each stage:

[0048]

[0049] Step Six: Information Matching: When the player presses the piano key X1, a signal f1 corresponding to the pressed piano key will be sent from the piano side. The signal f1 corresponds to the piano key X1 frame. Detect whether there is a human hand on the piano keyboard through Step Four. If so, detect the key points of the human hand through Step Five, and compare whether each key point near the fingertip falls within the piano key frame calibrated by X1. If there are calibrated fingertip key points within the frame, the finger used by the player can be determined through the fingertip key points, and then it can be identified whether the playing fingering is correct. If it is incorrect, an error prompt will be given.

Claims

1. An intelligent method for identifying piano fingering, characterized in that, it includes the following steps: Step 1: Camera installation and debugging: Install a camera whose viewing angle can fully cover the piano keyboard, and the camera can be connected to the piano display screen to display the piano keyboard image on the piano screen in real time; Step 2: Piano keyboard segmentation: Use the saliency object detection algorithm SOD100K based on deep learning for piano keyboard segmentation. The lightweight network proposed by this algorithm includes a feature extractor and a cross-stage fusion part, which can process features of multiple scales simultaneously; The feature extractor is stacked with the intra-layer multi-scale blocks proposed by SOD100K and is divided into 4 stages according to the resolution of the feature map. Each stage has 3, 4, 6, and 4 intra-layer multi-scale blocks respectively; The cross-stage fusion part composed of a flexible convolutional module (gOctConvs) proposed by SOD100K will process the features from each stage of the feature extractor to obtain a high-resolution output; Step 3: Piano keyboard calibration: Through the frame coordinates of the segmented keyboard obtained in Step 2, sort the keyboard from left to right and calibrate the keyboard frames such as X1, X2,...; Step 4: Detection of the human hand playing the piano: Collect the finger video of the piano being played by the camera and collect human hand pictures on the Internet and label the human hand frames and left and right classification marks, and use the FaceBoxes detection algorithm to train the human hand detector model; Step 5: Detection of key points of the playing fingers: Use the OpenPose human hand key point detection algorithm to detect the coordinates of the key points of the playing fingers, and mark the key points close to the fingertips as h1, h2, h3,...; Step 6: Joint matching of the information of the playing fingers and the piano keyboard: When the player presses the piano key X1, the piano side will send out a signal f1 corresponding to the pressed piano key. The signal f1 corresponds to the frame of the piano key X1. Detect whether there is a human hand on the piano keyboard through Step 4. If there is, detect the key points of the human hand through Step 5, and compare whether each key point close to the fingertip falls within the frame of the piano key calibrated by X1. If there are calibrated fingertip key points within the frame, the finger used by the player can be determined through the fingertip key points, and then whether the playing fingering is correct can be identified. If it is incorrect, an error prompt will be given.

Citation Information

Patent Citations

  • Video-game controller assemblies for progressive control of actionable-objects displayed on touchscreens

    CA2837808A1

  • Manipulator underactuated driving structure applied to piano teaching and design method

    CN110394784A