Motion capture method, system, storage medium, and program product
By combining piano key motion data and hand image recognition technology, and coordinating sensors and image processing, the system achieves accurate analysis and real-time feedback of piano playing effects. This solves the problems of low recognition rate and high hardware requirements in existing technologies, improves the accuracy and reliability of data collection, and reduces the burden on teachers and students.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CHANGSHA BAISI INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2025-11-21
- Publication Date
- 2026-05-28
Smart Images

Figure CN2025136670_28052026_PF_FP_ABST
Abstract
Description
A motion capture method, system, storage medium, and program product.
[0001] Priority application
[0002] This application claims priority to Chinese Invention Patent Application No. [202411697818.3], filed on November 25, 2024, entitled "[A motion capture method, system, storage medium and program product]", which is incorporated herein by reference in its entirety. Technical Field
[0003] This invention relates to the field of intelligent analysis technology for piano performance, specifically to a motion capture method, system, storage medium, and program product. Background Technology
[0004] Currently, most piano teaching relies on in-person instruction from teachers. This method is limited by factors such as manpower, time, money, and teacher skill level, significantly increasing the difficulty of learning piano. While some intelligent piano teaching systems have been proposed in the current technology, they suffer from at least the following major drawbacks:
[0005] Most solutions employ feature comparison-based methods. First, a standard database of correct hand shapes is established using a mathematical model. Then, a prediction model is built to extract features from predicted images, and these features are compared with the standard database to determine whether the hand shape is incorrect. However, feature comparison methods suffer from poor robustness, high subjectivity, and low recognition rates.
[0006] When building predictive models, many methods use stereo or depth cameras to obtain 3D data, thereby constructing 3D models. Compared to 2D vision models, 3D models are computationally intensive, complex in design, have poor performance, and require high-performance hardware, necessitating high-performance computing chips. Both depth cameras and high-performance computing chips significantly increase costs.
[0007] Furthermore, existing technologies for collecting hand information include the use of wearable devices (e.g., data gloves), motion tracking technologies (e.g., miniature radar systems), or manual extraction of gesture data from images. However, in piano playing, the use of wearable devices may affect the dexterity of the arm (or fingers), motion tracking technologies are not very accurate in detecting subtle movements such as finger movement on the keys or pressing the keys, and manual extraction of gesture data is labor-intensive, highly specialized, and lacks satisfactory generalization ability and robustness.
[0008] Therefore, there is an urgent need for a data collection and analysis method that can accurately reflect the performance effect of the performer. Summary of the Invention
[0009] The purpose of this invention is to provide a motion capture method, system, storage medium, and program product that partially solves or alleviates the above-mentioned shortcomings in the prior art and can improve the accuracy of motion capture data processing.
[0010] To solve the aforementioned technical problems, the present invention specifically adopts the following technical solution:
[0011] A first aspect of the present invention is to provide a motion capture method, comprising the steps of:
[0012] S1, acquire multiple sets of motion data for at least one first piano key, the motion data including: force, and / or displacement, the motion data being associated with a first timestamp;
[0013] S2, when the motion data indicates that the first key is pressed between the first time t1 and the second time t2, multiple sets of first hand images of the object are collected between the third time t3 and the fourth time t4; wherein the third time t3 is earlier than the first time t1 and the fourth time t4 is later than the second time t2; the first hand images are associated with a second timestamp.
[0014] S3, perform a pressing action or a first hand with a pressing tendency based on the first hand image recognition;
[0015] S4, identify the first finger in the first hand that performs the pressing action or has the pressing tendency;
[0016] S5, identify the second piano key associated with the first finger;
[0017] S6, determine whether the identifier of the second key matches that of the first key. If so, generate at least one set of motion capture results between the first moment and the second moment. The motion capture results include: the first hand and the first finger used to perform the pressing action, the identifier of the key, and the corresponding motion data. The motion capture results are associated with a corresponding third timestamp.
[0018] In some embodiments, the steps further include:
[0019] Multiple sets of second hand images are acquired, and each second hand image is associated with a fourth timestamp;
[0020] Based on the multiple sets of second hand images, the hand that performs the pressing action or has a pressing tendency is identified;
[0021] Identify the hand performing the pressing action or the fingers exhibiting the pressing tendency;
[0022] Identify a third key associated with the finger, and desired data for the third key, the desired data including: pressing time, and / or pressing force.
[0023] Obtain the motion data corresponding to the third piano key and the fourth timestamp;
[0024] Determine whether the expected data and the motion data match; if not, correct the motion capture result.
[0025] In some embodiments, when the first finger corresponding to the first key cannot be identified, the method further includes the step of:
[0026] Determine whether the first hand and the second hand are obstructed based on the third hand image;
[0027] If not, then identify at least one second finger in the first hand that failed to generate the motion capture result;
[0028] Specifically, when only one second finger is acquired, the second finger is associated with the first piano key.
[0029] In some embodiments, when the result of determining whether the first hand and the second hand occlude each other based on the third hand image is yes, the method further includes the step of:
[0030] Acquire multiple sets of fourth hand images between the fifth time t5 and the sixth time t6; wherein the fifth time t5 is earlier than the third time t3, and the sixth time t6 is later than the fourth time t4; the fourth hand images are associated with a fifth timestamp;
[0031] Based on the fourth hand image, the first hand and the second hand are identified at first positions at at least two first time nodes, where the first time nodes are located between t5 and t3;
[0032] The second positions of the first hand and the second hand are predicted based on the positions of at least two first time nodes, where the second time node is located between t3 and t4.
[0033] Predict the third position of the first finger based on the second position;
[0034] Based on the third location, query multiple sets of associated motion data.
[0035] The present invention also provides a motion capture system, comprising:
[0036] The first motion data acquisition module is used to acquire multiple sets of motion data for at least one first piano key. The motion data includes force and / or displacement, and the motion data is associated with a first timestamp.
[0037] The first image acquisition module is used to acquire multiple sets of first hand images of the object between the third time t3 and the fourth time t4 when the first piano key is identified to be in a pressed state between the first time t1 and the second time t2 based on the motion data; wherein the third time t3 is earlier than the first time t1 and the fourth time t4 is later than the second time t2; the first hand images are associated with a second timestamp.
[0038] The first hand recognition module is used to recognize a first hand that performs a pressing action or has a pressing tendency based on the first hand image;
[0039] The first finger recognition module is used to identify the first finger in the first hand that performs the pressing action or has the pressing tendency;
[0040] The first key recognition module is used to recognize the second key associated with the first finger;
[0041] The piano key matching module is used to determine whether the identifier of the second piano key matches that of the first piano key. If so, it generates at least one set of motion capture results between the first time and the second time. The motion capture results include: the first hand and the first finger used to perform the pressing action, the identifier of the piano key, and the corresponding motion data. The motion capture results are associated with a corresponding third timestamp.
[0042] In some embodiments, it also includes:
[0043] The second image acquisition module is used to acquire multiple sets of second hand images, and the second hand images are associated with a fourth timestamp;
[0044] The second hand recognition module is used to recognize the hand that performs the pressing action or has a pressing tendency based on the multiple sets of second hand images;
[0045] The second finger recognition module is used to identify the fingers of the hand that are performing the pressing action or have the pressing tendency;
[0046] The second key module is used to identify the third key associated with the finger, and the expected data of the third key, the expected data including: pressing time, and / or pressing force.
[0047] The second motion data module is used to acquire the motion data corresponding to the third piano key and the fourth timestamp;
[0048] The first correction module is used to determine whether the expected data and the motion data match; if not, the motion capture result is corrected.
[0049] In some embodiments, it also includes:
[0050] The second correction module is used to determine whether the first hand and the second hand are occluded based on the third hand image when the first finger corresponding to the first key cannot be identified; if not, it identifies at least one second finger in the first hand that failed to generate the motion capture result; wherein, when only one second finger is acquired, the second finger is associated with the first key.
[0051] In some embodiments, it also includes:
[0052] The third correction module is used to acquire multiple sets of fourth hand images between the fifth time t5 and the sixth time t6 when the result of determining whether the first hand and the second hand are occluded based on the third hand image is yes; wherein the fifth time t5 is earlier than the third time t3, and the sixth time t6 is later than the fourth time t4; the fourth hand images are associated with a fifth timestamp; the first positions of the first hand and the second hand at at least two first time nodes are identified based on the fourth hand images, the first time nodes being located between t5 and t3; the second positions of the first hand and the second hand at a second time node are predicted based on the positions of the at least two first time nodes, the second time nodes being located between t3 and t4; the third position of the first finger is predicted based on the second position; and multiple sets of associated motion data are queried based on the third position.
[0053] The present invention also provides a computer-readable storage medium storing a program or instructions that, when executed, implement the method as described in any of the embodiments.
[0054] The present invention also provides a computer program product storing a program or instructions that, when executed, implement the method described in any of the embodiments.
[0055] Beneficial technical effects:
[0056] This invention provides a method for the refined acquisition and analysis of an object's playing technique through the coordinated use of sensor data (such as force / displacement data) and image data. Specifically, by performing dual matching of sensor data and image data in both temporal and spatial dimensions, and by employing a two-layer decomposition analysis of the image, including hand and finger differentiation, motion capture data can be generated with precise correlation to the fingers.
[0057] Furthermore, this reliable motion capture data can accurately analyze the fingering of the subject (such as dynamics, speed, or rhythm). For example, the motion capture data can be compared with preset professional performance data for automatic analysis, or the motion capture data can be fed back to the piano teacher to assist the teacher in evaluating the subject's learning indicators and reduce the teacher's teaching workload.
[0058] Furthermore, this invention can also classify and analyze data mis-collection (such as detecting force sensing signals but not associating them with fingers; or detecting fingers lingering on piano keys for a long time but not detecting force sensing signals) by using methods such as hand differentiation, finger differentiation, and motion data matching, and correct errors in a timely manner, so as to further improve the accuracy and reliability of motion capture data and make this invention more practical.
[0059] It is worth noting that this invention aims to provide students and piano teachers with an auxiliary data detection and analysis system through the precise collection and intelligent analysis of motion capture data. Specifically, the intelligent data collection and analysis can provide students with more timely fingering feedback to improve their learning autonomy. Simultaneously, the real-time data collection also provides piano teachers with an auxiliary teaching assessment tool, reducing their teaching pressure and minimizing the time and space constraints of piano instruction.
[0060] Furthermore, this intelligent data detection and analysis system can also comprehensively control the cost of piano learning and reduce the learning burden on students. Attached Figure Description
[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. The elements or parts in the drawings are not necessarily drawn to scale. Obviously, the drawings described below are some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0062] Figure 1 is a schematic diagram of the method flow in an exemplary embodiment of the present invention;
[0063] Figure 2 is a schematic diagram of the system module structure in an exemplary embodiment of the present invention;
[0064] Figure 3 is a schematic diagram of the time axis correlation of the collected data in an exemplary embodiment of the present invention;
[0065] Figure 4 is a schematic diagram of the module structure of a computing device in an exemplary embodiment of the present invention. Detailed Implementation
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0067] In this document, suffixes such as "module," "part," or "unit" used to denote elements are used only for the purpose of illustrative purposes and have no specific meaning in themselves. Therefore, "module," "part," or "unit" may be used interchangeably.
[0068] In this document, the terms "upper," "lower," "inner," "outer," "front," "rear," "one end," and "the other end," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the present invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0069] In this document, unless otherwise explicitly specified and limited, the terms "installed," "equipped with," "connected," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0070] In this document, "and / or" includes any and all combinations of one or more of the listed related items.
[0071] In this article, "multiple" means two or more, that is, it includes two, three, four, five, etc.
[0072] As used in this specification, the term "about" typically means + / -5% of the value, more typically + / -4%, more typically + / -3%, more typically + / -2%, even more typically + / -1%, and even more typically + / -0.5%.
[0073] In this specification, certain embodiments may be disclosed in a range-bound format. It should be understood that this "range-bound" description is merely for convenience and brevity and should not be construed as a rigid limitation on the disclosed range. Therefore, the description of a range should be considered as having specifically disclosed all possible subranges and the individual numerical values within those ranges. For example, a description of the range 1-6 should be considered as having specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., and the individual numbers within those ranges, such as 1, 2, 3, 4, 5, and 6. This rule applies regardless of the breadth of the range.
[0074] A piano generally consists of keys (usually white and black keys) and a metal soundboard. The range of an 88-key piano is generally from about 27.5 Hz to about 4186.01 Hz, while that of a 108-key piano is up to about 7902.13 Hz.
[0075] Due to the specialized nature of piano playing, piano learners often rely on in-person instruction from professional piano teachers to practice and develop piano skills under their guidance. Furthermore, this dependence on in-person instruction makes teaching highly susceptible to geographical limitations; for example, when teachers and students are in different locations, it can be difficult for the teacher to accurately assess the student's progress.
[0076] Accordingly, referring to Figures 1-4, this invention provides a motion capture method based on the synergy of sensor data and image data. This method enables precise digital analysis of a user's (e.g., a student's) playing performance using more accurate and reliable motion capture data. This method provides users with more direct and accurate data feedback while significantly assisting piano teachers in completing professional teaching tasks. Furthermore, this motion capture method combining sensor data and image data reduces the technical difficulty of data processing while improving data reliability.
[0077] Example 1
[0078] Referring to Figure 1, the present invention provides a motion capture method, including the following steps:
[0079] S1, acquire multiple sets of motion data for at least one first piano key, the motion data including: force, and / or displacement, the motion data being associated with a first timestamp;
[0080] Specifically, the force here can be the pressure exerted on the piano key, which is used to provide feedback on the user's finger pressure; the displacement can be the pressing depth of the piano key, or the rotation angle of the piano key at a certain time.
[0081] For example, in some embodiments, a speed sensor, displacement sensor or acceleration sensor is provided below the piano keys, which can detect the speed, displacement or acceleration of the corresponding piano keys in real time, and thus indirectly calculate the pressing depth of the piano keys or the force applied.
[0082] S2, when the motion data indicates that the first key is pressed between the first time t1 and the second time t2, multiple sets of first hand images of the object are collected between the third time t3 and the fourth time t4; wherein the third time t3 is earlier than the first time t1 and the fourth time t4 is later than the second time t2; the first hand images are associated with a second timestamp.
[0083] For example, if the pressing depth of the first key is greater than a set depth threshold, then the first key is considered to be in a pressed state.
[0084] S3, perform a pressing action or a first hand with a pressing tendency based on the first hand image recognition;
[0085] In some embodiments, at least one image acquisition module (such as a camera) is used to acquire an image of the hand of an object (or user).
[0086] In some embodiments, the camera may include any suitable device capable of acquiring images and converting them into electrical signals. The camera may include, but is not limited to, a video photometer, an infrared scanner, a video camera, etc.
[0087] In some embodiments, one or more cameras may be configured to capture images of the hand. For example, in some embodiments, a first camera may be configured to capture lateral images of the hand, capable of identifying the degree of finger flexion. As another example, in some embodiments, a second camera may be configured above the piano keys to identify the degree of pressure applied by each finger.
[0088] In some embodiments, the collected data may also be one or more video data consisting of multiple sets of hand images.
[0089] S4, identify the first finger in the first hand that performs the pressing action or has the pressing tendency;
[0090] When an object plays a piece of music, the processor can acquire hand image information from the camera, recognize the hand image, extract user gesture information, that is, recognize the hand (such as distinguishing between the left and right hands), and then recognize the pressing action or the first finger with the pressing tendency based on the degree of bending and pressing of each finger in the hand.
[0091] In some embodiments, when the degree of bending of a finger is large, it can be preliminarily determined that the finger has made a pressing action; or, in some embodiments, when the degree of bending of a finger is relatively small, it can be determined that it may produce or perform a pressing action, and therefore it is preliminarily determined that the finger has a pressing tendency.
[0092] In some embodiments, when multiple players are involved in a game, hand features (such as hand size, skin texture, wrist joint features, etc.) can be used to initially identify two (or even multiple) objects. That is, the objects are first associated with hand images, and then left and right hands and fingers are distinguished.
[0093] S5, identify the second piano key associated with the first finger;
[0094] In some embodiments, the camera can calculate or infer the actual position of the first finger at the current moment from the hand image. Matching the actual position with the position of the piano keys can identify the second key actually operated by the first finger, such as by obtaining the identifier of the second key. The identifier can be a key number (ID), name, or other data used to identify the key.
[0095] For example, in some embodiments, the information about the piano keys is predetermined, such as the coordinates of the keys or the relative positions of multiple keys, or the positional relationships of the keys can be recorded using images (e.g., taking one or more images including parts or all of the keys). Correspondingly, when it is necessary to locate the key played by a finger, the position information can be determined by comparing the position of the key observed in the image with the predetermined position, or the real-time image observed by the camera can be compared with the recorded image to identify the key.
[0096] S6, determine whether the identifier of the second key matches that of the first key. If so, generate at least one set of motion capture results between the first moment and the second moment. The motion capture results include: the first hand and the first finger used to perform the pressing action, the identifier of the key, and the corresponding motion data. The motion capture results are associated with a corresponding third timestamp.
[0097] As an exemplary implementation, this embodiment preferably uses sensors as the primary means of acquiring precise data, while using images as an auxiliary means of predicting the degree of pressing posture (such as pressing action or pressing trend). By combining precise acquisition with degree prediction, more reliable motion capture results are output.
[0098] It is worth noting that, taking the auxiliary teaching scenario as an example, the motion capture method of this invention can analyze the playing state of a piano student (i.e., the object) in real time (e.g., whether the fingering is accurate). However, real-time analysis places extremely high demands on computing power. To address this, this embodiment uses sensors for precise data acquisition, and images are used in conjunction with posture-assisted verification. This improves the accuracy and reliability of motion capture data while reducing the performance requirements for real-time analysis to a certain extent through relatively less image processing.
[0099] Therefore, from another perspective, this embodiment can reduce the implementation cost of motion capture methods to a certain extent.
[0100] In some embodiments, the present invention also provides a data verification method, the method comprising the steps of:
[0101] Multiple sets of second hand images are acquired, and each second hand image is associated with a fourth timestamp;
[0102] Based on the multiple sets of second hand images, the hand that performs the pressing action or has a pressing tendency (or referred to as the third hand) is identified;
[0103] Identify the hand performing the pressing action or a finger (or referred to as a third finger) that has the pressing tendency;
[0104] Identify the third piano key associated with the finger, and the expected data for the third piano key, the expected data including: pressing time (or expected time), and / or pressing force (or expected force); for example, in some embodiments, the expected data can be preset by the user based on information from the piano score. Obtain the motion data of the third piano key corresponding to the fourth timestamp;
[0105] Determine whether the expected data and the motion data match;
[0106] If so, the motion capture result can be generated and output accordingly;
[0107] If not, the motion capture results can be corrected. For example, if the difference between the expected data and the motion data falls within a set range, they are considered to match; otherwise, they are considered to not match.
[0108] In other words, this embodiment also provides an error identification and correction function.
[0109] For example, in some embodiments, for a specific piano score, each note in the score can typically have common performance standards, such as performance duration, performance dynamics, etc., through which expected data can be preset. Alternatively, in some embodiments, expected data refers to a reasonable range of key pressing time or pressure intensity predicted based on pressing actions or pressing trends identified in hand images.
[0110] For example, if hand image analysis indicates that a finger has pressed or tended to press a piano key (i.e., the third key) for an extended period, but the expected data is not obtained for that key (e.g., no force signal is captured, or only a brief signal is captured, significantly differing from the prediction result of the hand image), then it is considered that there may be an error in data acquisition or analysis, requiring correction of the data processing procedure. For example, in some embodiments, the correction process may involve updating the hand or finger recognition results using more hand images.
[0111] For example, in some embodiments, the correction may be to correct the finger numbering or piano key numbering of the motion capture results.
[0112] For example, in some embodiments, when the first finger corresponding to the first key cannot be identified, the method further includes the step of:
[0113] Determine whether the first hand and / or the second hand are obstructed based on the third hand image;
[0114] If not, then identify at least one second finger in the first hand that failed to generate the motion capture result;
[0115] Specifically, when only one second finger is acquired, the second finger is associated with the first piano key.
[0116] In some embodiments, due to the subject employing different types of playing techniques (e.g., finger crossing or finger straddling), or perhaps due to errors made by the subject during playing (e.g., finger bending, wrist concave, wrist too high, and metacarpophalangeal joint collapse), it may result in partial obstruction of at least one hand (or part of the fingers) of the subject from the camera's perspective.
[0117] In some embodiments, the third hand image can be any image between the third time t3 and the fourth time t4.
[0118] In some embodiments, the third hand image can also be any image between the first time t1 and the second time t2.
[0119] In this embodiment, when it is detected that the first key is pressed by an object, but no corresponding finger (i.e., the first finger) is associated with it, an elimination method can be used to query for potentially associated fingers. Specifically, in this embodiment, when there are unassociated fingers in the object's left hand, the system first uses image retrieval to determine whether the object's left and right hands are occluded. If not occluded, it checks whether the other fingers of the left hand have been successfully matched and generated corresponding motion capture results. If so, it queries for fingers that have not successfully generated motion capture results. When there is a finger that has not successfully generated motion capture results, it is recommended to associate that finger with the first key.
[0120] This invention proposes a motion capture data association scheme that is primarily sensor-based and secondarily image-based. Specifically, this invention assigns different acquisition tasks to sensor acquisition and image acquisition: sensors are used to accurately acquire force (which is indirectly calculated through position, velocity, or acceleration), while images are used to acquire posture information of the hand. Furthermore, this invention performs secondary matching on the data sequentially from the time dimension and the piano key dimension to quickly complete the data association.
[0121] It is worth noting that this multi-coordinated approach, which integrates data collection and matching methods, can achieve a good balance between data collection reliability and data processing pressure.
[0122] In some embodiments, when the result of determining whether the first hand and the second hand occlude each other based on the third hand image is yes, the method includes the steps of:
[0123] Multiple sets of fourth hand images are acquired between the fifth time t5 and the sixth time t6; wherein the fifth time t5 is earlier than the third time t3, and the sixth time t6 is later than the fourth time t4; each fourth hand image is associated with a fifth timestamp.
[0124] Based on the fourth hand image, the first hand and the second hand are identified at first positions at at least two first time nodes, where the first time nodes are located between t5 and t3;
[0125] The first hand and the second hand are predicted at a second time node based on the first position of at least two first time nodes, where the second time node is located between t3 and t4.
[0126] For example, in some embodiments, the movement speed of the hand can be calculated based on multiple hand images (such as a fourth hand image). (For example, the movement speed can refer to the speed at which the hand moves in a planar direction.) The position of the hand at the next time point can be predicted based on the movement speed.
[0127] For example, in some embodiments, the movement trend of the hand can be calculated based on multiple hand images, such as predicting the position at the next moment based on the movement trend at the current moment.
[0128] Furthermore, in some embodiments, a third position of the first finger is predicted based on the second position; then, multiple sets of associated motion data are queried based on the third position.
[0129] Alternatively, in some embodiments, the third position of the first finger is predicted based on the second position, and if the third position corresponds to the first key, a motion capture result is generated based on the motion data of the first key and the first finger.
[0130] In some embodiments, "association" means that the piano key associated with the motion data matches the piano key corresponding to the third position. For example, in some embodiments, if multiple sets of associated motion data are found, these multiple sets of motion data can be associated with the first finger to form a motion capture result.
[0131] For example, based on the position of the hand and the relative position of the fingers to the hand (such as by identifying the fingers as the thumb or index finger, the approximate position of the fingers in the hand can be determined), the predicted position of the fingers (i.e., the third position) can be obtained.
[0132] This embodiment describes another error correction method when the hand is occluded. For example, when the left and right hands of the object are partially occluded, the position of the hand during the occlusion period is predicted using at least one segment of image data before and after the occlusion. Furthermore, the position of the fingers can be predicted based on the image data before and after the occlusion. The prediction result is then matched with the actual detected sensor data (i.e., motion data). If the match is successful, the current result is used to complete the correction.
[0133] To more clearly describe the technical solution adopted in this invention, Figure 3 provides an exemplary depiction of the motion data and the time axis of the first and fourth hand images in this embodiment.
[0134] In other words, in this embodiment, in order to reconcile the contradiction between the requirements of real-time data output and accuracy of motion capture results, it is preferable to take the sensor as the lead, that is, when the sensor cannot directly match a valid finger, the situation of possible occlusion is verified.
[0135] It should be noted that performing highly complex pieces often involves intricate hand gestures, such as crossed hands or four-handed duets, which significantly increases the difficulty of real-time motion capture analysis. Therefore, in this embodiment, when sensor data cannot be directly matched, an occlusion recognition process is preferably enabled to reduce the computational burden and difficulty of real-time analysis while still completing the task.
[0136] Furthermore, by designing a collaborative recognition mode that prioritizes sensors and uses image processing as an auxiliary tool, the real-time processing range of images can be expanded without excessively increasing the burden on image processing. For example, it can perform relatively complete recognition of multiple images with a wider temporal field of view (such as covering multiple moments before, during, and after the keys are pressed) in a short period of time, which also helps to improve the reliability of images as an auxiliary verification method. This avoids the difficulty in complete, clear, or reliable finger recognition caused by excessively fast hand speeds during complex playing techniques.
[0137] Alternatively, from another perspective, this embodiment preferably prioritizes the time result generated by the sensor (such as the timestamp of the moment of pressing), selecting the temporal field of view (or timeline) for image processing based on the sensor's time. This dual coordination of the time dimension and the functional task dimension can effectively balance the accuracy and computational burden of real-time analysis.
[0138] In some embodiments, at least one sensor is provided under at least one piano key, and the sensor is configured to detect one or more of the following data: displacement, velocity, and acceleration of the piano key.
[0139] In some embodiments, the sensor may be a velocity sensor, a displacement sensor, or an acceleration sensor.
[0140] Furthermore, in some embodiments, the speed sensor includes, but is not limited to, one or more of the following: magnetoelectric speed sensor, Hall effect speed sensor, photoelectric speed sensor, and magnetoelectric speed sensor.
[0141] Furthermore, in some embodiments, the displacement sensor includes, but is not limited to, one or more of the following: potentiometer-type displacement sensor, inductive displacement sensor, capacitive displacement sensor, magnetostrictive displacement sensor, optical sensor, ultrasonic sensor, Hall effect displacement sensor, oscillator-type displacement sensor, photoelectric encoder, contact sensor, and non-contact sensor, such as eddy current displacement sensor, laser displacement sensor, ultrasonic displacement sensor, etc., which do not directly contact the object being measured.
[0142] For example, in other embodiments, sound sensors may also be provided on the piano keys, which can sense the signals generated by pressing the keys to determine which key was pressed.
[0143] It is worth noting that this invention can be flexibly applied to multiple application scenarios:
[0144] For example, when a piano teacher needs to conduct remote online teaching for students, the motion capture data can be fed back to the teacher's port in real time, which can help the teacher complete the teaching task in a remote location.
[0145] For example, when students practice after class, they can use motion capture data to correct their own mistakes. Alternatively, parents can use the self-analysis function of motion capture data to supervise their children's practice process.
[0146] Example 2
[0147] Referring to Figure 2, the present invention also provides a motion capture system, comprising:
[0148] The first motion data acquisition module 10 is used to acquire multiple sets of motion data for at least one first piano key. The motion data includes force and / or displacement, and the motion data is associated with a first timestamp.
[0149] The first image acquisition module 20 is used to acquire multiple sets of first hand images of the object between the third time t3 and the fourth time t4 when the first piano key is identified to be in a pressed state between the first time t1 and the second time t2 based on the motion data; wherein the third time t3 is earlier than the first time t1 and the fourth time t4 is later than the second time t2; the first hand images are associated with a second timestamp.
[0150] The first hand recognition module 30 is used to recognize a first hand that performs a pressing action or has a pressing tendency based on the first hand image;
[0151] The first finger recognition module 40 is used to recognize the first finger in the first hand that performs the pressing action or has the pressing tendency.
[0152] The first key recognition module 50 is used to recognize the second key associated with the first finger;
[0153] The piano key matching module 60 is used to determine whether the identifier of the second piano key matches that of the first piano key. If so, it generates at least one set of motion capture results between the first time and the second time. The motion capture results include: the first hand and the first finger used to perform the pressing action, the identifier of the piano key, and the corresponding motion data. The motion capture results are associated with a corresponding third timestamp.
[0154] In some embodiments, it also includes:
[0155] The second image acquisition module is used to acquire multiple sets of second hand images, and the second hand images are associated with a fourth timestamp;
[0156] The second hand recognition module is used to recognize the hand that performs the pressing action or has a pressing tendency based on the multiple sets of second hand images;
[0157] The second finger recognition module is used to identify the fingers of the hand that are performing the pressing action or have the pressing tendency;
[0158] The second key module is used to identify the third key associated with the finger, and the expected data of the third key, the expected data including: pressing time, and / or pressing force.
[0159] The second motion data module is used to acquire the motion data corresponding to the third piano key and the fourth timestamp;
[0160] The first correction module is used to determine whether the expected data and the motion data match; if not, the motion capture result is corrected.
[0161] In some embodiments, it also includes:
[0162] The second correction module is used to determine whether the first hand and the second hand are occluded based on the third hand image when the first finger corresponding to the first key cannot be identified; if not, it identifies at least one second finger in the first hand that failed to generate the motion capture result; wherein, when only one second finger is acquired, the second finger is associated with the first key.
[0163] In some embodiments, it also includes:
[0164] The third correction module is used to acquire multiple sets of fourth hand images between the fifth time t5 and the sixth time t6 when the result of determining whether the first hand and the second hand are occluded based on the third hand image is yes; wherein the fifth time t5 is earlier than the third time t3, and the sixth time t6 is later than the fourth time t4; the fourth hand images are associated with a fifth timestamp; the first positions of the first hand and the second hand at at least two first time nodes are identified based on the fourth hand images, the first time nodes being located between t5 and t3; the second positions of the first hand and the second hand at a second time node are predicted based on the positions of the at least two first time nodes, the second time nodes being located between t3 and t4; the third position of the first finger is predicted based on the second position; and multiple sets of associated motion data are queried based on the third position.
[0165] The present invention also provides a computer-readable storage medium storing a program or instructions that, when executed, implement the method as described in any of the embodiments.
[0166] The present invention also provides a computer program product storing a program or instructions that, when executed, implement the method described in any of the embodiments.
[0167] Furthermore, the present invention also provides an electronic device, comprising:
[0168] One or more processors;
[0169] One or more memory units;
[0170] And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, and the one or more computer programs include instructions that, when executed by the one or more processors, cause the electronic device to perform the method as described in any embodiment.
[0171] Furthermore, the present invention also provides a computer-readable storage medium storing a program or instructions that, when executed, implement the method described in any embodiment.
[0172] Furthermore, the present invention also provides a computer program product storing a program or instructions that, when executed, implement the method described in any embodiment.
[0173] Referring to Figure 4, the present invention provides a computing device, and the computer device provided by the present invention includes a processor, a memory and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and internal memory.
[0174] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any motion capture method.
[0175] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0176] Internal memory provides an environment for the execution of computer programs stored in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to perform any motion capture method.
[0177] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that the structure shown in Figure 4 is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. Specific computer devices may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0178] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0179] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a computer terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0180] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A motion capture method characterized by, Including the following steps: S1, acquire multiple sets of motion data for at least one first piano key, the motion data including: force, and / or displacement, the motion data being associated with a first timestamp; S2, when the motion data indicates that the first key is pressed between the first time t1 and the second time t2, multiple sets of first hand images of the object are collected between the third time t3 and the fourth time t4; wherein the third time t3 is earlier than the first time t1 and the fourth time t4 is later than the second time t2; the first hand images are associated with a second timestamp. S3, perform a pressing action or a first hand with a pressing tendency based on the first hand image recognition; S4, identify the first finger in the first hand that performs the pressing action or has the pressing tendency; S5, identify the second piano key associated with the first finger; S6, determine whether the identifier of the second key matches that of the first key. If so, generate at least one set of motion capture results between the first moment and the second moment. The motion capture results include: the first hand and the first finger used to perform the pressing action, the identifier of the key, and the corresponding motion data. The motion capture results are associated with a corresponding third timestamp.
2. The motion capture method of claim 1, wherein, It also includes the following steps: Multiple sets of second hand images are acquired, and each second hand image is associated with a fourth timestamp; The pressing action or a third hand with the pressing trend is executed based on the recognition of the multiple sets of second hand images; Identify the third hand performing the pressing action or the third finger having the pressing tendency; Identify the third key associated with the third finger, and the desired data for the third key, the desired data including: pressing time, and / or pressing force; Obtain the motion data corresponding to the third piano key and the fourth timestamp; Determine whether the expected data and the motion data match; if not, correct the corresponding motion capture result.
3. The motion capture method of claim 1, wherein, When the first finger corresponding to the first key cannot be identified, the following steps are also included: Determine whether the first hand and the second hand are obstructed based on the third hand image; If not, then identify at least one second finger in the first hand that failed to generate the motion capture result; Specifically, when only one second finger is acquired, the second finger is associated with the first piano key.
4. The motion capture method of claim 3, wherein, If the result of determining whether the first hand and the second hand are occluded based on the third hand image is yes, then the following step is also included: Multiple sets of fourth hand images are acquired between the fifth time t5 and the sixth time t6; wherein the fifth time t5 is earlier than the third time t3, and the sixth time t6 is later than the fourth time t4; each fourth hand image is associated with a fifth timestamp. Based on the fourth hand image, the first hand and the second hand are identified at first positions at at least two first time nodes, the first time nodes being located between the fifth time t5 and the third time t3; The first hand and the second hand are predicted at a second time node based on the first position of at least two first time nodes, where the second time node is located between the third time t3 and the fourth time t4; Predict the third position of the first finger based on the second position; Based on the third location, query multiple sets of associated motion data.
5. A motion capture system characterized by, include: The first motion data acquisition module is used to acquire multiple sets of motion data for at least one first piano key. The motion data includes force and / or displacement, and the motion data is associated with a first timestamp. The first image acquisition module is used to acquire multiple sets of first hand images of the object between the third time t3 and the fourth time t4 when the first piano key is identified to be in a pressed state between the first time t1 and the second time t2 based on the motion data; wherein the third time t3 is earlier than the first time t1 and the fourth time t4 is later than the second time t2; the first hand images are associated with a second timestamp. The first hand recognition module is used to recognize a first hand that performs a pressing action or has a pressing tendency based on the first hand image; The first finger recognition module is used to identify the first finger in the first hand that performs the pressing action or has the pressing tendency; The first key recognition module is used to recognize the second key associated with the first finger; The piano key matching module is used to determine whether the identifier of the second piano key matches that of the first piano key. If so, it generates at least one set of motion capture results between the first time and the second time. The motion capture results include: the first hand and the first finger used to perform the pressing action, the identifier of the piano key, and the corresponding motion data. The motion capture results are associated with a corresponding third timestamp.
6. The motion capture system of claim 5, wherein, Also includes: The second image acquisition module is used to acquire multiple sets of second hand images, and the second hand images are associated with a fourth timestamp; The second hand recognition module is used to recognize the pressing action or a third hand with the pressing trend based on the multiple sets of second hand images; The second finger recognition module is used to recognize the third finger that performs the pressing action or has the pressing tendency; The second key module is used to identify the third key associated with the third finger, and the expected data of the third key, the expected data including: pressing time, and / or pressing force. The second motion data module is used to acquire the motion data corresponding to the third piano key and the fourth timestamp; The first correction module is used to determine whether the expected data and the motion data match; if not, the motion capture result is corrected.
7. The motion capture system of claim 5, wherein, Also includes: The second correction module is used to determine whether the first hand and the second hand are occluded based on the third hand image when the first finger corresponding to the first key cannot be identified; if not, it identifies at least one second finger in the first hand that failed to generate the motion capture result; wherein, when only one second finger is acquired, the second finger is associated with the first key.
8. The motion capture system of claim 7, wherein, Also includes: The third correction module is used to acquire multiple sets of fourth hand images between the fifth time t5 and the sixth time t6 when the result of determining whether the first hand and the second hand are occluded based on the third hand image is yes; wherein the fifth time t5 is earlier than the third time t3, and the sixth time t6 is later than the fourth time t4; the fourth hand images are associated with a fifth timestamp; the first positions of the first hand and the second hand at at least two first time nodes are identified based on the fourth hand images, the first time nodes being located between the fifth time t5 and the third time t3; the second positions of the first hand and the second hand at a second time node are predicted based on the first positions of the at least two first time nodes, the second time nodes being located between the third time t3 and the fourth time t4; the third position of the first finger is predicted based on the second position; and multiple sets of associated motion data are queried based on the third position.
9. A computer-readable storage medium, characterized in that, The storage medium stores a program or instructions that, when executed, implement the method as described in any one of claims 1 to 4.
10. A computer program product, characterised in that, The computer program product stores a program or instructions that, when executed, implement the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Piano with scoring function and scoring method thereof
CN106935226A
Detection device of piano playing intonation
CN107103895A
Piano teaching method, device and computer storage medium
CN109215441A
Piano playing strength teaching aid device
CN110379255A
Intelligent auxiliary method and device for piano practice and medium
CN113657185A