A fingering recognition method, system, device and medium applied to a musical instrument system

By setting up an scalable time window and attention adjustment mechanism in the musical instrument system to dynamically match musical instrument performance events, the problems of accuracy and computational resource waste in musical instrument fingering recognition are solved, resulting in more efficient fingering recognition and a better user experience.

CN121459429BActive Publication Date: 2026-05-29WUXIAN HONGYIN (CHONGQING) TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUXIAN HONGYIN (CHONGQING) TECHNOLOGY CO LTD
Filing Date
2026-01-06
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-precision fingering recognition in complex musical instrument performance scenarios, and they also result in significant waste of computing resources.

Method used

An image acquisition module is used to set an expandable time window, dynamically adjust attention based on the triggering and releasing events of the driving element, merge time windows of highly correlated hand movements, and optimize the allocation of computing resources.

Benefits of technology

It improves the accuracy of finger recognition and the efficiency of computing resource utilization, adapts to the complex scenarios of musical instrument performance, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459429B_ABST
    Figure CN121459429B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of musical instrument performance recognition, in particular to a fingering recognition method, system, device and medium applied to a musical instrument system, the musical instrument system comprising a musical instrument, the musical instrument comprising a driving element, a sensor and an image acquisition module, the image acquisition module comprising at least two image acquisition units arranged towards the musical instrument, the at least two image acquisition units being used for acquiring at least two images, and the at least two images having different visual angles; correspondingly, the method comprises the steps that a first time point when a triggering event of the driving element occurs and a second time point when release events of the driving element occur in sequence are acquired; a window updating model is used to set corresponding time windows for the triggering event through the first time point and the second time point; wherein the window updating model comprises a time window = [the first time point - lambda 1, the second time point + lambda 2]. The application can improve the fingering recognition precision and reduce the operation pressure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of musical instrument performance recognition technology, specifically to a fingering recognition method, system, device, and medium applied to musical instrument systems. Background Technology

[0002] In recent years, the number of self-taught musical instrument users has surged, and the technology for automatic musical instrument performance recognition has also flourished accordingly.

[0003] For example, patent application CN114783222A provides a musical instrument-assisted teaching method and system based on AR technology, including a data acquisition step, in which a teaching video is played through an AR device, and the actual video data and actual audio data during the performance are acquired simultaneously; a judgment step, in which the loudness peak of the pitch in the actual audio data is acquired to determine the actual rhythm point, the audio at the actual rhythm point is acquired to determine whether a playing error has occurred, and the actual fingering image in the actual video data is extracted from the actual rhythm point to determine whether a fingering error has occurred; and an error correction step, in which the playing error or fingering error at the actual rhythm point is acquired, and the correct audio data or correct fingering image corresponding to the actual rhythm point is retrieved, and the teaching is guided by voice through audio comparison or image comparison.

[0004] For example, patent application CN116912935A discloses an action recognition method, system, and storage medium. The method includes: detecting the action of a performer's fingers based on a preset deep learning algorithm to generate multiple hand detection boxes corresponding to the fingers; classifying the type of fingers according to the hand detection boxes to obtain the performer's finger position information, wherein the finger position information is used to characterize the movement of the performer's fingers during performance; acquiring the positions of preset piano keys to obtain reference keys; identifying the joints of the performer's fingers based on a preset posture recognition algorithm to obtain a set of finger joints; calculating the distance between each element in the set of finger joints and the reference keys to determine the piano note corresponding to the finger position information; and determining the target finger action based on the finger position information and the piano note.

[0005] For example, patent application CN120183041A provides an intelligent music teaching assistance system and method based on motion perception. It can use deep learning algorithms based on artificial intelligence and image analysis to automatically analyze students' playing movements, compare the differences between the movements and the standard musical playing movements of the instrument, and perform temporal analysis to identify problems in the students' movements. It also provides targeted feedback and improvement suggestions on the continuity of the students' playing movements, thereby assisting in intelligent music teaching.

[0006] While the above methods can automate the recognition or evaluation of users' playing actions to a certain extent, they cannot well adapt to the complex scenarios and refined requirements of musical instrument playing action recognition. There is an urgent need for a finger recognition method that can better fit the field of musical instrument playing recognition technology. Summary of the Invention

[0007] The purpose of this invention is to provide a finger recognition method, system, device and medium for musical instrument systems, which partially solves or alleviates the above-mentioned shortcomings in the prior art, improves finger recognition accuracy and reduces computational pressure.

[0008] To solve the aforementioned technical problems, the present invention specifically adopts the following technical solution:

[0009] A first aspect of the present invention is to provide a fingering recognition method applied to a musical instrument system, the musical instrument system comprising: a musical instrument, the musical instrument comprising: a driving element, which, when manipulated, can directly or indirectly control the musical instrument to produce sound; a sensor for detecting the manipulation state of the driving element, the manipulation state including: triggering and releasing; and an image acquisition module comprising: at least two image acquisition units disposed facing the musical instrument, the at least two image acquisition units being used to acquire at least two images, and the at least two images having different viewing angles;

[0010] Correspondingly, the method includes the following steps:

[0011] S101, determine whether a trigger event has occurred in the driving element;

[0012] If S101 is true, then S102 is executed:

[0013] S102, Collect the first moment when the triggering event occurs in the driving element, and the second moment when the release event occurs in sequence in the driving element;

[0014] S103, a window update model is used to set a corresponding time window for the triggering event based on the first time point and the second time point; wherein, the window update model includes:

[0015] The time window = [first moment - λ1, second moment + λ2]; where λ1 is the set first extended time and λ2 is the set second extended time.

[0016] S104, within the time window, the attention of the image acquisition module is increased or maintained at the first level of attention;

[0017] The image acquisition module is used to generate a finger matching result based on the at least two images. The finger matching result includes the matching relationship between at least one finger and the driving element.

[0018] In some embodiments, the steps further include:

[0019] S105, update the attention to a second attention level within an interval window; wherein the interval window is a time period during which the triggering event has not occurred; the second attention level is less than the first attention level.

[0020] In some embodiments, the interval window is the time period between two adjacent time windows.

[0021] In some embodiments, the musical instrument is a piano, and the driving element is a piano key.

[0022] In some embodiments, prior to S104, the following step is also included:

[0023] When at least two adjacent trigger events are identified, the positions of the corresponding at least two driving elements and the corresponding at least two time windows are obtained;

[0024] When the positional interval between at least two of the driving elements is less than a preset spacing threshold, the at least two time windows are merged into one time window.

[0025] In some embodiments, the steps further include:

[0026] When adjacent release events and trigger events are identified sequentially, the adjacent difference between the times corresponding to the release event and the trigger event is calculated;

[0027] When the adjacent difference is greater than a preset time interval, the first extension time is set to the first extension value, and / or the second extension time is set to the second extension value;

[0028] When the adjacent difference is less than the time interval, the first extended time is set to the third extended value, and / or the second extended time is set to the fourth extended value;

[0029] Wherein, the first extended value is less than the third extended value, and the second extended value is less than the fourth extended value.

[0030] A second aspect of the present invention is to provide a fingering recognition system for a musical instrument system, the musical instrument system comprising: a musical instrument, the musical instrument comprising: a driving element, which, when manipulated, can directly or indirectly control the musical instrument to produce sound; a sensor for detecting the manipulation state of the driving element, the manipulation state including: triggering and releasing; and an image acquisition module comprising: at least two image acquisition units disposed facing the musical instrument, the at least two image acquisition units being used to acquire at least two images, and the at least two images having different viewing angles;

[0031] Correspondingly, the system includes:

[0032] A trigger judgment module is used to determine whether a trigger event has occurred in the driving element. When the judgment result of the trigger judgment module is yes, the time acquisition module is entered: the time acquisition module is used to acquire the first moment when the driving element triggers the event and the second moment when the driving element releases the event in sequence; a window setting module is used to set a corresponding time window for the trigger event using a window update model based on the first moment and the second moment; wherein, the window update model includes: the time window = [first moment - λ1, second moment + λ2]; wherein, λ1 is the set first extension time and λ2 is the set second extension time; an attention update module is used to increase or maintain the attention of the image acquisition module to a first attention level under the time window; wherein, the image acquisition module is used to generate a finger matching result based on the at least two images, and the finger matching result includes: the matching relationship between at least one finger and the driving element.

[0033] In some embodiments, the system further includes:

[0034] An interval window module is used to update the attention to a second attention level within an interval window; wherein the interval window is a time period during which the triggering event has not occurred; and the second attention level is less than the first attention level.

[0035] A third aspect of the present invention is to provide a computer device, the device including a memory and a processor; the memory being used to store a computer program; the processor being used to execute the computer program and, when executing the computer program, to implement a fingering recognition method for a musical instrument system as described in any embodiment of the present invention.

[0036] A fourth aspect of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement a fingering recognition method for a musical instrument system as described in any embodiment of the present invention.

[0037] Beneficial technical effects:

[0038] This invention proposes a fingering recognition method for musical instrument systems. It can apply higher attention to more critical fingering movements at relatively appropriate times, thereby selectively focusing on key segments within complex fingering data, improving recognition accuracy, and to some extent achieving dynamic temporal matching between analysis resources and performance events, reducing computational burden. Specifically:

[0039] 1) In view of the fast and complex characteristics of playing musical instruments, this invention provides an attention update mechanism for an image acquisition module, that is, a higher first level of attention is applied during the time window when the driving element is triggered, and a lower second level of attention is applied during the interval window when it is not triggered, thereby largely avoiding the waste of resources caused by attention, while ensuring the accuracy of finger recognition.

[0040] 2) For scenarios involving continuous key presses within a short period, this invention preferably merges and analyzes hand movements with close intervals between piano keys. Firstly, by merging highly correlated time windows, it effectively avoids repeatedly switching attention states within multiple adjacent time windows in a short period (state switching may incur additional computational resource overhead). Secondly, it can merge at least two isolated time windows into a more complete time window (which can reflect highly correlated triggering events and the hand shape changes between them). This not only enables more coherent and complete recognition of the user's finger movements but also helps the model learn the hand shape changes between triggering events, helping users solidify correct hand shapes and finger usage habits.

[0041] 3) By leveraging the difference between adjacent release events and subsequent trigger events, time windows of different lengths can be adaptively set, making finger recognition, skill analysis, and subsequent feedback guidance more in line with the actual situation of musical expression, greatly enhancing user experience and application value. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. The elements or parts in the drawings are not necessarily drawn to scale. Obviously, the drawings described below are some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0043] Figure 1 This is a flowchart illustrating a fingering recognition method applied to a musical instrument system, as provided in an embodiment of this application.

[0044] Figure 2 This is a flowchart illustrating an event recognition method applied to a musical instrument system, as provided in an embodiment of this application.

[0045] Figure 3 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application;

[0046] Figure 4 This is a schematic diagram of the structure of a fingering recognition system applied to a musical instrument system, provided in an embodiment of this application;

[0047] Figure 5 This is a schematic diagram of the structure of an event recognition system applied to a musical instrument system, provided in an embodiment of this application. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0049] In this document, suffixes such as "module," "component," or "unit" used to denote elements are used solely for the purpose of illustrative purposes and have no specific meaning in themselves. Therefore, "module," "component," or "unit" can be used interchangeably. In this document, terms such as "upper," "lower," "inner," "outer," "front," "rear," "one end," and "the other end," indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In this document, unless otherwise expressly specified and limited, terms such as "installed," "equipped with," and "connected" should be interpreted broadly. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium; it can be a connection within two elements. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. In this document, "and / or" includes any and all combinations of one or more of the listed related items. "A plurality of" means two or more, i.e., it includes two, three, four, five, etc. As used in this specification, the term "about" typically means + / -5% of the value, more typically + / -4%, more typically + / -3%, more typically + / -2%, even more typically + / -1%, even more typically + / -0.5%. In this specification, certain embodiments may be disclosed in a range format. It should be understood that this "range" description is merely for convenience and brevity and should not be construed as a rigid limitation on the disclosed range. Therefore, the description of the range should be considered as having specifically disclosed all possible subranges and independent numerical values ​​within those ranges. For example, a description of the range 1-6 should be considered as having specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as individual numbers within this range, such as 1, 2, 3, 4, 5, and 6. The above rules apply regardless of the breadth of the range.

[0050] Definition of the noun:

[0051] In this article, the musical instrument can have keys arranged in a fixed order, and each key has a fixed pitch. These keys can be combined to form a keyboard, and the driving element can be the keys themselves. The musical instrument can be a piano, organ, accordion, or electronic keyboard, etc., and is not limited here.

[0052] In this article, the image acquisition module is a device that converts optical signals in the physical world into digital images, such as cameras and camcorders. Its core components may include photoelectric converters (such as CCD or CMOS sensors), lenses, signal processing units, and storage modules. The model and specific components of the image acquisition module can be selected according to actual needs.

[0053] For example, the installation location of the image acquisition module can be determined based on specific application scenarios (such as ambient light) and technical parameters (such as resolution, installation height, angle, and image integrity) to ensure device performance and compliance. For instance, the image acquisition module can be perpendicular to the piano key plane, located directly above the keyboard, such as inside the piano lid, on the top of the piano stand, or on a hanging bracket; or, for another example, the image acquisition module can be installed at a certain angle to the piano key plane, on the front or upper side of the piano. Alternatively, the image acquisition module can be independently positioned corresponding to the user's viewing angle, or it can be configured as a head-mounted device that changes its viewing angle according to the user's perspective.

[0054] Example 1:

[0055] When a user plays an instrument, a camera (equivalent to an image acquisition unit) can be used to track and recognize their hand movements to identify whether their specific fingering (i.e., the way the user's fingers move when playing, such as in the piano field, fingering includes fingering, finger crossing, finger extension, finger contraction, etc.) is standard, and then their playing methods and playing habits can be corrected in a targeted manner.

[0056] In traditional musical instrument learning and intelligent assistance methods, finger recognition is a key indicator for evaluating learning effectiveness. To achieve accurate judgment of the relationship between fingers and keys, traditional solutions rely on continuous high frame rate image capture and uninterrupted real-time computing and video analysis.

[0057] However, many problems arise when these traditional finger recognition schemes are applied to real, dynamic instrument playing scenarios.

[0058] For example, real instrumental performance is highly dynamic, meaning that fingering changes in real time depending on factors such as the note arrangement, tempo, tonality, and techniques in the score. In other words, the rapid and complex finger movements during performance place higher demands on the real-time response and precise capture capabilities of fingering recognition.

[0059] In response to the fingering recognition scenario in the aforementioned musical instrument performance, this invention sets an expandable time window for fingering recognition, which can apply higher attention to more critical fingering movements at a relatively appropriate time. This allows for selective focus on key segments in complex fingering data, thereby achieving dynamic matching of analysis resources and performance events in terms of time sequence to a certain extent.

[0060] In some embodiments, see Figure 1 This invention provides a fingering recognition method for a musical instrument system. The musical instrument system includes: a musical instrument, the musical instrument including: a driving element, which, when manipulated, can directly or indirectly control the musical instrument to produce sound; a sensor for detecting the manipulation state of the driving element, the manipulation state including: triggering and releasing; and an image acquisition module, the image acquisition module including: at least two image acquisition units arranged facing the musical instrument, the at least two image acquisition units for acquiring at least two images, and the at least two images having different viewing angles.

[0061] Correspondingly, the method includes the following steps:

[0062] S101, determine whether a trigger event has occurred in the driving element;

[0063] If S101 is true, then S102 is executed:

[0064] S102, Collect the first moment when the triggering event occurs in the driving element, and the second moment when the release event occurs in sequence in the driving element;

[0065] S103, a window update model is used to set a corresponding time window for the triggering event based on the first time point and the second time point; wherein, the window update model includes:

[0066] The time window = [first moment - λ1, second moment + λ2]; where λ1 is the set first extended time and λ2 is the set second extended time.

[0067] S104, within the time window, the attention is increased or maintained at the first level of attention;

[0068] The image acquisition module is used to generate a finger matching result based on the at least two images. The finger matching result includes the matching relationship between at least one finger and the driving element.

[0069] The following will describe the method and steps for generating finger matching results with reference to specific embodiments:

[0070] 1. System initialization: Calibrate the two image acquisition units (such as parameters, installation angle, etc.) and set the sensors.

[0071] 2. Real-time acquisition: Synchronously acquire sensor data and images.

[0072] 3. Image preprocessing: Correct distortion and extract the ROI (Region of Interest) of the hand.

[0073] 4. Hand detection and key point localization: Hand key points are detected in the images from the left and right image acquisition units respectively.

[0074] 5. Keypoint Fusion: Transform the keypoints from two perspectives to the same coordinate system and perform a weighted average or select points with high confidence.

[0075] 6. Finger recognition: Based on the fused key points, the position of each fingertip is calculated. Combined with the key press information provided by the sensor, the finger corresponding to each pressed key is determined through nearest neighbor matching or geometric relationship matching.

[0076] 7. Image Fusion and Annotation: Fusion of images from two perspectives into a top-down view image, drawing key points of the fingers on the image, and labeling the pressed keys with finger numbers.

[0077] 8. Display and Feedback: Display the merged image and finger placement information to the user and provide learning feedback.

[0078] The sensors may include pressure and / or capacitive sensors to accurately collect piano key trigger information.

[0079] Among them, the image resolution is preferably above 1080P, and the image frame rate is preferably above 60fps.

[0080] Among them, the SPI (Serial Peripheral Interface) high-speed master-slave communication protocol can be used to enable multiple hardware devices (such as cameras and sensors) to operate on the same reference clock and generate high-precision timestamps to achieve time alignment.

[0081] In some embodiments, a virtual top view can be generated by performing perspective transformation on at least two images.

[0082] It should be understood that by setting an expandable time window, the present invention helps to capture other key actions besides the user's key touch actions, thereby greatly improving the accuracy of finger recognition while saving computing resources.

[0083] For example, the first extended time can be retrospectively observed before the chord is pressed, including the preparatory movements of the fingers involved. For instance, the degree of hand opening and the relative positions of the 1st and 5th fingers before playing an octave. This can greatly improve the robustness of chord fingering recognition.

[0084] For example, the second extended time can continuously observe whether the finger slides to the next key or is replaced by another finger on the original key after it is lifted, thereby enabling a more complete analysis of the overall finger movement and obtaining more accurate results.

[0085] In some embodiments, the musical instrument is a piano, and the driving element is a piano key.

[0086] In some embodiments, the actuator may be a user's body part (such as a finger) and / or an automated assistive device (such as a robotic arm).

[0087] In some embodiments, the sound production principle of a piano is as follows: when a user presses a key, the key drives the hammer to strike the corresponding string quickly through the internal action mechanism (including levers, push rods, and other components). After being struck, the string vibrates at a specific frequency. The vibration sound waves are amplified and resonated through the soundbox (or soundboard) inside the piano body, ultimately forming the piano sound. When the key is released, the damper resets and presses down on the string, the vibration stops, and the piano sound stops.

[0088] In some embodiments, the foot pedal can change the state of the piano's dampers, thereby altering the timbre and duration of the sound that is subsequently or is being produced.

[0089] In other words, when the driving element is a piano key, the key can directly control the instrument to produce sound; when the driving element is a piano pedal, the pedal can indirectly control the timbre or duration of the instrument's sound.

[0090] In some embodiments, the sensor detects the control state of the driving element in the following ways: a speed sensor, displacement sensor or acceleration sensor can be provided below the piano keys, which can detect in real time the force (which can be calculated based on speed, displacement or acceleration), speed, displacement or acceleration of the corresponding key being triggered.

[0091] For example, when the force of triggering the piano key is greater than or equal to the preset trigger force, and / or the speed is greater than or equal to the preset trigger speed, and / or the displacement is greater than or equal to the preset trigger displacement, and / or the acceleration is greater than or equal to the preset trigger acceleration, the control state of the piano key can be considered as triggered.

[0092] The preset trigger force, and / or preset trigger speed, and / or preset trigger displacement, and / or preset trigger acceleration are typically a very small threshold. As long as this threshold is exceeded (e.g., the preset trigger displacement can be 10% of the maximum displacement of the driving element), the driving element can be considered to be triggered.

[0093] In some embodiments, when a key is triggered (or pressed), if the sensor detects that the pressure and / or displacement begins to decrease, the key can be considered to be in a released state.

[0094] The (trigger) displacement refers to the amount of movement of a driving element (such as a piano key) from its initial rest position to a certain position.

[0095] In some embodiments, the image acquisition module is used to simultaneously acquire image sequences and / or video sequences from at least two perspectives through image acquisition units (such as cameras) with different perspectives, thereby providing a more reliable data source for subsequent finger recognition.

[0096] In some embodiments, setting at least two image acquisition units to acquire at least two images or videos has at least the following advantages: 1) avoiding finger occlusion leading to distortion of finger recognition; 2) cross-validation of multi-source data to calculate more accurate three-dimensional coordinates of finger joints; 3) with the help of at least two cameras, a wider field of view can be covered, while ensuring to a certain extent that the control state of each driving element can be acquired relatively completely.

[0097] In some embodiments, the at least two image acquisition units may be positioned above or to the side front of the instrument to create a cross-view.

[0098] Taking a piano with two image acquisition units as an example, one camera is placed in the upper left of the keyboard, facing the lower right; the other is placed in the upper right, facing the lower left. The fields of view of the two cameras can overlap significantly in the piano key area.

[0099] Alternatively, one camera might be set up to provide an overview of the overall positional relationships and movement trends of all fingers and actuating elements; another camera might be positioned behind or diagonally behind the keys to capture the vertical displacement (depth direction) of finger presses, as well as details below the fingertips that might be obscured by the back of the hand from the main viewpoint.

[0100] In some embodiments, the field of view of at least two image acquisition units should be adjusted to fully cover all driving elements (such as piano keys) and achieve maximum overlap within the space where the performer's hands move.

[0101] The space for the performer's hand movements can be determined based on the performer's historical performance data. For example, for performers who prefer exaggerated performance movements, the space for hand movements can be set relatively large.

[0102] In some embodiments, at least two image acquisition units should be precisely aligned in time when they are working to avoid errors in the calculation of the three-dimensional coordinates of the finger joints due to delays in the two acquired images.

[0103] In some embodiments, the effects of piano key reflections or ambient light on image quality can be avoided by adjusting the angle of the image acquisition unit, adding a polarizing filter, or supplementing with a controllable light source.

[0104] In some embodiments, if a trigger event occurs in the driving element, the first moment when the trigger event occurs and the second moment when the release event occurs can be collected, and a window update model can be used to set a corresponding time window for a complete trigger (at least including the key touch process from trigger to release).

[0105] In some embodiments, the window update model can define the duration of the time window's forward and backward expansion based on λ1 and λ2, thereby updating the time window from the actual keystroke duration to an observation duration that better suits the needs of real performance.

[0106] In some embodiments, λ1 and / or λ2 may also be 0.

[0107] In some embodiments, λ1 and λ2 can be used to characterize the length of the extension time; that is, the larger λ1 and λ2 are, the longer the extension time, and correspondingly, the longer the time window.

[0108] λ1 is used to define the time backward from the moment the driving element is triggered. Its core purpose is to analyze the user's hand movements a certain period of time before the driving element is actually pressed.

[0109] λ2 is used to define the time extended after the moment the drive element is released. Its core purpose is to continue the analysis of the user's hand movements for a period of time after the drive element is released.

[0110] In some embodiments, the at least two images may be at least two images (or pictures).

[0111] In some embodiments, the at least two images may be at least two sequences of moving images (or video clips).

[0112] In some embodiments, finger matching results can also be generated based on at least two images in the following manner:

[0113] 1) Capture image sequences from the at least two cameras within the same time window.

[0114] 2) In each frame of the image sequence, use a computer vision model (such as a deep learning-based instance segmentation model) to identify the various parts of the hand (including fingertips, knuckles, back of the hand, etc.) and the corresponding driving elements (such as piano keys that may be and are actually triggered and the edges of the keys).

[0115] 3) Locate the two-dimensional pixel coordinates of each part of the hand and the two-dimensional pixel coordinates of the key trigger point (such as the leading edge of the key).

[0116] 4) Based on the two-dimensional pixel coordinates from different cameras at the same time, use a stereo vision algorithm (based on pre-calibrated camera parameters, such as perspective angle) to calculate the precise coordinates (X, Y, Z) of each two-dimensional pixel coordinate in three-dimensional space.

[0117] 5) Connect the three-dimensional coordinates of the key hand points in each frame to construct a dynamic three-dimensional hand skeleton model to simulate the user's real hand playing movements.

[0118] In some embodiments, the relationship between the three-dimensional coordinates of each fingertip and the three-dimensional coordinates of the corresponding driving element can be continuously analyzed and determined. Then, combined with the timestamp of the driving element being triggered by the sensor, it can be determined which finger triggered which driving element at what time. For example, whether the Z-axis coordinate (height) of the fingertip and the triggered key is lower than the Z-axis coordinate of the key surface (i.e., whether the key has been pressed down), and whether its X and Y coordinates fall within the boundary range of a specific key.

[0119] In some embodiments, the finger identifier (such as the right ring finger) with the same timestamp and the corresponding drive element identifier (such as the center C key), as well as the control state (such as trigger or release), can be output as a set of finger matching results.

[0120] It should be understood that image acquisition modules with different specifications can generate different fingering matching results. For example, within the same time window, the higher the acquisition frequency (i.e., the specifications) of the image acquisition module, the more fingering matching results will be output. The specifications of the image acquisition module can be configured according to the user's actual performance needs, and are not limited here.

[0121] In some embodiments, the control state also includes holding down, i.e., after triggering and before releasing, the user holds the pressing action for a duration exceeding a set duration.

[0122] In some embodiments, the musical instrument may be a stringed instrument (such as a guitar or guzheng), and the driving element may be a string.

[0123] For example, in some embodiments, the sensor may refer to a device that determines the pressing / movement state of a driving element (such as a string) based on an image. For example, the sensor may include a camera, and the camera may share a connection with an image acquisition module, meaning the camera can be used simultaneously to perform finger placement and driving element recognition tasks.

[0124] In some embodiments, the steps further include:

[0125] S105, update the attention of the image acquisition module to a second attention level within an interval window; wherein, the interval window is a time period during which the triggering event has not occurred; the second attention level is less than the first attention level.

[0126] In some embodiments, under the interval window, it can be assumed that the user is not triggering the driving element at this time. At this time, a second level of attention can be applied to the interval window (such as reducing the image acquisition frequency, reducing the image resolution, etc.) to save the computing resources, storage bandwidth and energy consumption of finger matching.

[0127] The applicant noted that playing musical instruments (especially technical expression) often requires users to change fingering movements at high speeds, and it is very difficult to continuously recognize fingering movements under high load during long periods of practice or performance.

[0128] Meanwhile, the performance is event-driven, meaning that the start and end of each note are explicitly dependent on the user's performance action (such as triggering or releasing). When the user does not trigger the driving element, it can be considered that this is an idle period with low analytical value.

[0129] In other words, unlike continuous monitoring, this invention provides an attention update mechanism for the image acquisition module, taking into account the rapid and complex nature of musical instrument performance. Specifically, a higher level of primary attention is applied during the time window when the driving element is triggered, while a lower level of secondary attention is applied during the interval window when it is not triggered. This largely avoids resource waste caused by attention issues while ensuring the accuracy of finger placement recognition.

[0130] In some embodiments, the interval window is the time period between two adjacent time windows.

[0131] In some embodiments, prior to S104, the following step is also included:

[0132] When at least two adjacent trigger events are identified, the positions of the corresponding at least two driving elements and the corresponding at least two time windows are obtained;

[0133] When the positional interval between at least two of the driving elements is less than a preset spacing threshold, the at least two time windows are merged into one time window.

[0134] "Adjacent" can refer to the time interval between at least two triggering events being less than a preset time interval (such as 1 second).

[0135] "Adjacent" can refer to two consecutive triggering events occurring in adjacent times, such as the first trigger and the second trigger; or it can refer to two non-consecutive triggering events occurring in adjacent times, such as the first trigger and the third trigger.

[0136] In some embodiments, for scenarios involving continuous key touches within a short period, the present invention preferably performs combined analysis on hand movements with closely spaced key positions. If the user is monitored rapidly jumping between closely spaced keys (such as chords, dense clusters of notes, or rapid consecutive notes), the context of the gesture change is often highly correlated. In this case, continuous recognition can be performed on the video, that is, at least two time windows can be merged into one time window to improve recognition accuracy through more continuous recognition.

[0137] Furthermore, this invention does not identify finger A pressing key X and finger B pressing key Y in isolation, but rather identifies the global hand posture adopted by the user to perform rapid, continuous operations (such as pressing keys X and Y consecutively). Merging time windows helps to analyze at least two triggering events and their preceding and following hand movements as a complete posture change process in a coherent manner, while also optimizing attention allocation efficiency with the help of coherent time windows.

[0138] It should be understood that this merging analysis mechanism has at least the following technical effects: 1) By merging highly correlated time windows, it effectively avoids repeatedly switching attention levels in multiple adjacent time windows within a short period (state switching may cause additional computational resource overhead). 2) Merging at least two isolated time windows into a more complete time window (which can reflect highly correlated triggering events and the hand shape change process between them) not only enables more coherent and complete recognition of the user's finger movements, but also helps the model learn the hand shape change movements between triggering events, helping users solidify correct hand shapes and finger usage habits.

[0139] In some embodiments, the steps further include:

[0140] When adjacent release events and trigger events are identified sequentially, the adjacent difference between the times corresponding to the release event and the trigger event is calculated;

[0141] When the adjacent difference is greater than a preset time interval, the first extension time is set to the first extension value, and / or the second extension time is set to the second extension value;

[0142] When the adjacent difference is less than the time interval, the first extended time is set to the third extended value, and / or the second extended time is set to the fourth extended value;

[0143] Wherein, the first extended value is less than the third extended value, and the second extended value is less than the fourth extended value.

[0144] In some embodiments, if λ1 and λ2 are 0, then the adjacent difference between the release event and the trigger event is equal to the time length of the corresponding interval window.

[0145] In some embodiments, if the time interval between the previous release event and the subsequent trigger event is relatively long, it may indicate that the user's finger movements (such as lifting and lowering) are skilled and usually have clear movement characteristics. In this case, the observation time window can be appropriately shortened (i.e., using shorter first and second extension values). If the time interval between the previous release event and the subsequent trigger event is very short (such as fast passages or vibrato), it may indicate that the user's playing habit is to play close to the keys (i.e., the fingers do not need to be deliberately lifted, but only the fingertips touch the keys with a small force from the knuckles or metacarpophalangeal joints, avoiding large movements of the arm or wrist). In this case, the trend of hand movement changes may not be very obvious, and the model cannot identify when and how the user's fingers move to play the next note. Therefore, the observation time window can be appropriately extended (i.e., using longer third and fourth extension values) to capture more hand movements that are not perceptible in a short time (such as a slight lifting of the hand before triggering), thereby improving the accuracy of recognition.

[0146] For example, in some embodiments, if there is a long time interval between the previous release event and the next trigger event, it may indicate that the user is raising their finger high (such as a clear hand raising-lowering process). In this case, the observation time window can be appropriately shortened (i.e., a longer first extension value and second extension value are used) to avoid misjudging the deliberate pause as a continuous action, while significantly reducing the average power consumption.

[0147] It should be understood that by using the magnitude of the adjacent difference between the previous release event and the subsequent trigger event, time windows of different lengths can be adaptively set, making finger recognition, skill analysis, and subsequent feedback guidance more in line with the actual situation of musical expression, greatly improving the user experience and application value.

[0148] In some embodiments, see Figure 4The present invention also provides a fingering recognition system applied to a musical instrument system, the musical instrument system comprising: a musical instrument, the musical instrument comprising: a driving element, which, when manipulated, can directly or indirectly control the musical instrument to produce sound; a sensor, the sensor being used to detect the manipulation state of the driving element, the manipulation state including: triggering and releasing; an image acquisition module, the image acquisition module comprising: at least two image acquisition units positioned facing the musical instrument, the at least two image acquisition units being used to acquire at least two images, and the at least two images having different perspectives; correspondingly, the system comprising: a trigger judgment module, used to determine whether the driving element has triggered an event; when the determination result of the trigger judgment module is yes, then entering a time acquisition module: the time acquisition module being used to acquire the first moment when the driving element triggers the event, and the second moment when the driving element releases the event; a window setting module, used to set a corresponding time window for the trigger event using a window update model based on the first moment and the second moment; wherein, the window update model includes: the time window = [First moment - λ1, Second moment + λ2]; where λ1 is the set first extension time and λ2 is the set second extension time; Attention update module, used to raise or maintain the attention of the image acquisition module to the first attention level under the time window; wherein, the image acquisition module is used to generate finger matching results based on the at least two images, the finger matching results including: the matching relationship between at least one finger and the driving element.

[0149] In some embodiments, the system further includes:

[0150] An interval window module is used to update the attention to a second attention level within an interval window; wherein the interval window is a time period during which the triggering event has not occurred; and the second attention level is less than the first attention level.

[0151] It should be understood that the finger recognition system applied to the musical instrument system can be used to implement the method steps described in any embodiment of the present invention.

[0152] In some embodiments, this application also provides a schematic block diagram of the structure of a computer device, please see... Figure 3 Computer programs can be used in situations such as Figure 3 It runs on the computer device shown. Figure 3As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The memory may include non-volatile storage media and internal memory. The non-volatile storage media may store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor to execute any fingering recognition method applied to a musical instrument system. The processor provides computational and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the execution of the computer program in the non-volatile storage media; when executed by the processor, this program causes the processor to execute any fingering recognition method applied to a musical instrument system. The network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 3 The structures shown are merely block diagrams of a portion of the structure related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. It should be understood that the processor may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0153] Example 2:

[0154] The applicant noted that the individual differences in instrumental performance are very prominent, with significant differences in the precision, speed, and stability of hand movements among performers of different skill levels.

[0155] Especially for fast and complex musical passages, traditional fingering recognition methods may miss key fingering details and waste computational resources. In addition, continuous recognition requires processing more data, placing higher demands on hardware and real-time performance. Furthermore, if invalid segments (such as still hands) are interspersed in continuous playing, noise must be filtered out using algorithms.

[0156] In response, this invention proposes an attention update mechanism based on the fit between the sheet music and the user (e.g., the relationship between the performance difficulty level and the user level threshold), which can greatly improve the fingering recognition effect in different performance scenarios and for different user groups, while taking into account both recognition accuracy and computational resource utilization efficiency.

[0157] In some embodiments, see Figure 2 This invention proposes an event recognition method for a musical instrument system, the musical instrument system comprising: a musical instrument, the musical instrument comprising: a driving element, which, when manipulated, can directly or indirectly control the musical instrument to produce sound; a sensor, the sensor being used to detect the manipulation state of the driving element, the manipulation state including: triggering and releasing; and an image acquisition module, the image acquisition module comprising: at least two image acquisition units positioned facing the musical instrument, the at least two image acquisition units being used to acquire at least two images, and the at least two images having different perspectives; wherein, the image acquisition module is used to generate a finger matching result based on the at least two images, the finger matching result including: the matching relationship between at least one finger and the driving element;

[0158] Correspondingly, the method includes the following steps:

[0159] S201, determine whether a trigger event has occurred in the driving element;

[0160] If S201 is true, then S202 is executed:

[0161] S202, Collect the first moment when the driving element triggers an event, and the second moment when the driving element releases an event sequentially;

[0162] S203, a window update model is used to set a corresponding time window for the triggering event based on the first time point and the second time point; wherein, the window update model includes:

[0163] The time window = [first moment - λ1, second moment + λ2]; where λ1 is the set first extended time and λ2 is the set second extended time.

[0164] S204, within the time window, the attention of the image acquisition module is increased or maintained at the first level of attention;

[0165] The updating or setting steps of the window update model include:

[0166] S205, Get the performance type of the performance;

[0167] S206, Select a recommended extension time from the first extension time and the second extension time according to the performance type;

[0168] In some embodiments, the method further includes the step of updating the window update model according to at least one of the recommended extended times.

[0169] In this embodiment, a more personalized recommended extension time is selected based on the performance type. In other words, the extension of the time window can be tailored to different performance types, such as extending only forward or only backward. This can improve the finger recognition effect while saving computing resources.

[0170] In some embodiments, the emphasis may be placed on extending the time window before and after different performance types (such as performance techniques).

[0171] In some embodiments, users can preset the recommended extension time based on a type of performance (such as performance technique).

[0172] In some embodiments, the method of selecting recommended extended time from the extended time according to different performance types can be set according to the actual performance needs of the performance type and / or the user's practice preferences.

[0173] For example, selecting the first extension time and the second extension time together as the recommended extension time is generally suitable for playing strong chords, accents, high notes of the melody, and phrase initiation notes.

[0174] For example, selecting only the first extension time as the recommended extension time is generally applicable to the starting note of ornaments (tremolo, appoggiatura), the ending note of a large leap, and the playing of a note that needs to be emphasized in a fast passage.

[0175] For example, selecting only the second extension time as the recommended extension time is generally suitable for the release of sustained notes, long notes, starting notes of large leaps, and the playing of notes in legato phrases.

[0176] In some embodiments, the performance type includes: performance position, performance technique and / or melody type, and the performance type is marked with at least one of the recommended extended times.

[0177] Among them, the playing position refers to the relative position of the currently played note within the musical phrase.

[0178] For example, if the performance position is close to the beginning or end of a musical phrase, a relatively long time window can be set to capture the details of the transition between phrases (such as the preparatory movements before the beginning of a phrase or the closing movements after the end of a phrase).

[0179] Furthermore, if the playing position is near the beginning of a musical phrase, the first extension time can be preferred as the recommended extension time to capture the preparatory movements before the beginning of the phrase; if the playing position is near the end of a musical phrase, the second extension time can be preferred as the recommended extension time to capture the closing movements after the end of the phrase, thereby enabling a more comprehensive analysis of the fingering movements.

[0180] In some embodiments, a musical phrase can be a minimum segmental unit that can be repeatedly trained by dividing the score according to bar lines (vertical lines).

[0181] Among them, performance techniques (including vibrato, glissando, legato, etc.) refer to the performance patterns adopted to produce specific musical effects, which can correspond to specific hand movement patterns.

[0182] In some embodiments, when complex techniques are involved, advance preparation is necessary. For example, when playing the starting note of a musical phrase, the high note of a melody, or a complex chord, the preceding and following notes can be expanded simultaneously. Taking accented playing as an example, before applying force with the hand, one can observe whether the user's arm is raised, whether the wrist is ready, and whether the hand shape is pre-positioned. After release, one can observe the hand movements to infer whether the user's hand strength has been effectively controlled and relaxed, thereby assessing whether the user's force application method is correct and whether incorrect force application will lead to long-term muscle tension, and thus correcting the user's incorrect playing habits.

[0183] For example, when dealing with passing notes in fast passages (such as scales and arpeggios), the core goal for the performer is to achieve fluency, evenness, and clarity, rather than assigning unique expression or color to each individual passing note. The key to playing these notes lies in the overall evenness of the continuous translational movement of the fingers, rather than the subtle gestures of individual fingers touching the keys. From an efficiency-first perspective, this can be avoided by not extending the time window, or by only extending it forward, as long as the key hand points are not lost. This avoids excessive analysis of individual passing notes, which could negatively impact the model's recognition of the user's overall hand movements and key notes.

[0184] For example, playing techniques may include vibrato, glissando, and legato. Vibrato is an ornamental playing technique, also known as embellishment, which involves rapidly and evenly alternating between a note and the note a second above (or occasionally below) it, creating a continuous undulating effect. Glissando refers to the rapid, continuous sliding of the fingers along the keys, allowing a series of adjacent pitches (usually white keys or notes arranged in a specific key) to be sounded smoothly and seamlessly in succession, creating a smooth, flowing, or brilliant auditory effect through continuous pitch transitions. Legato refers to a playing method where adjacent notes are sounded without noticeable gaps, maintaining a smooth transition in timbre and dynamics through continuous finger touches and natural connections. It is usually marked with an arc above or below the note, and its core is to achieve seamless connection between notes through techniques such as advance finger movement and touch-the-key, creating a lyrical and coherent melodic line.

[0185] Melody type can refer to the style of the musical score.

[0186] For example, the melody type can be uplifting, exciting, cheerful, gloomy, or tense.

[0187] For example, different melodic types can be preset with different recommended extension times. For instance, for gentler melodic types, a second extension time is preferred as the recommended extension time to improve the capture of transitional movements. Conversely, for more energetic melodic types, a first extension time is preferred as the recommended extension time to analyze the accuracy of the preparatory movements for hand exertion.

[0188] In some embodiments, a mapping table of different performance types and different recommended extension times may be pre-developed by music experts, performers and / or engineers.

[0189] In some embodiments, obtaining the performance type includes the following steps:

[0190] Obtain the sheet music data for the performance;

[0191] The section to be played in the next time period is determined based on the musical score data;

[0192] Identify the performance type of the paragraph.

[0193] In some embodiments, the performance type can be determined in advance through score data analysis.

[0194] In some embodiments, the sheet music data may be a digital sheet music file that the user is playing, loaded before or during the performance, such as MusicXML (Music Extensible Markup Language), MIDI (Musical Instrument Digital Interface), etc.

[0195] In some embodiments, the performance progress can be monitored in real time, and the position of the score can be accurately located by comparing the actual triggered keys with the note sequence in the score. At the same time, the score content of the next segment to be played can be previewed to identify at least one performance type of the segment to be played in the next segment.

[0196] For example, when complex techniques are involved (such as the starting note of a musical phrase, the high note of a melody, or complex chords), the extension time can be increased in advance to prioritize the allocation of computing resources to extend the time window before and after, ensuring that key fingerings are accurately captured.

[0197] The duration of the next time period can be preset (e.g., 10 seconds).

[0198] In some embodiments, the duration of the next time period may be the duration of the next musical segment.

[0199] In some embodiments, the performance technique for the next period can be identified based on the technique markers in the score (such as legato, staccato, accent, vibrato, arpeggio, etc.).

[0200] For example, if the passage to be played in the next time period contains a vibrato marker, then the playing technique is vibrato. Furthermore, the corresponding recommended extension time can be selected based on the vibrato playing technique.

[0201] In some embodiments, users may preset at least one of the following as the performance type: performance position, performance technique, and / or melody type.

[0202] It should be understood that this invention, through in-depth analysis of musical score data, proposes a restrictive recommended extension time selection scheme, which can significantly improve the accuracy and efficiency of fingering recognition from the perspective of adapting to the musical score. This restrictive recommended extension time selection scheme is particularly suitable for complex and fast passages. Specifically, different recommended extension time selection schemes are determined for different performance types, and the computational resources for fingering analysis are adaptively increased or decreased. This ensures that no key fingering information is missed, and also simplifies non-key information, allowing attention to be focused on the true technical points of instrument performance.

[0203] In some embodiments, the steps further include:

[0204] Obtain the user's historical performance data;

[0205] A performance score is generated based on the historical performance data.

[0206] The window update model is updated based on the performance score.

[0207] In some embodiments, historical performance data may refer to some or all of the historical performance data accumulated by the current user; alternatively, historical performance data may refer to recent performance-related data generated within a specified time period (e.g., 10 minutes) prior to the current time point. The specific settings are left to the user and are not limited here.

[0208] In some embodiments, a performance scoring model can be trained to score historical performance data.

[0209] The performance scoring model may include:

[0210] S301, acquire user data and historical performance data from several group dimensions. The historical performance data includes driving element information, the first moment, and the degree of triggering (including the force, speed, displacement, acceleration, etc. of the triggering).

[0211] Among them, the driving element information can refer to the number of the specific piano key that is triggered.

[0212] Specifically, individualized user data (such as playing ability) can be obtained through user pre-input or external device collection, and fingering data such as the first moment of triggering the driving element, the second moment of releasing the driving element, and the degree of triggering can be obtained in real time through sensors during the performance.

[0213] S302, during the user's performance, a first score is determined based on the first moment / second moment and the corresponding expected moment, and / or the driving element and the corresponding expected element; a second score is determined based on the actual triggering degree of the driving element and the expected triggering degree.

[0214] Specifically, the system analyzes the current musical score in real time, determines the expected timing based on the ideal trigger times specified by the notes in the score, and identifies the expected target based on the driving element corresponding to the notes in the score. The first score quantifies timing accuracy by comparing the user's actual trigger / release timing with the expected trigger / release timing, and / or the consistency between the actual driving element and the expected element, judging whether the user "played correctly." The smaller the deviation, the higher the first score.

[0215] In some embodiments, the performance score may be a first score.

[0216] Actual triggering level refers to the physical response generated during user triggering, as measured by sensors, including one or more parameters such as trigger force, speed, displacement, and acceleration. Correspondingly, the current musical score is analyzed in real time, and standard values ​​for trigger force, displacement depth, and speed are determined based on the standard triggering intensity or depth requirements of different musical styles, rhythmic sections, or the current position on the score, thus obtaining the expected triggering level. The second score quantifies the physical execution quality based on the degree of matching between the actual and expected triggering levels, judging whether the user "played well." The smaller the difference between the two, the higher the second score.

[0217] In some embodiments, the performance score may be a second score.

[0218] Alternatively, in some embodiments, the performance score may be the sum of a first score and a second score.

[0219] It should be understood that performance scores can reflect a user's fingering habits to some extent, such as whether fingering movements conform to fingering standards. Adjusting the observation time window for user fingering movements based on different performance scores can greatly improve the adaptability of this invention in the field of musical instrument performance movement recognition technology, ensuring that recognition accuracy is improved while saving computational resources.

[0220] In some embodiments, updating the window update model based on the performance score includes the following steps:

[0221] When the performance score is greater than the preset score threshold, the first extended time is set to the first extended value, and / or the second extended time is set to the second extended value;

[0222] When the performance score is less than the score threshold, the first extended time is set to the third extended value, and / or the second extended time is set to the fourth extended value;

[0223] Wherein, the first extended value is less than the third extended value, and the second extended value is less than the fourth extended value.

[0224] In other words, the present invention can update the extension time of the time window according to the actual performance of the user. When the performance score is high, it can indicate that the user's performance effect is better (or the user has a higher level of mastery of the current performance content, and the corresponding triggering and release actions will be more skillful and neat). At this time, the first extension time and the second extension time can be set to smaller first extension value and second extension value respectively, that is, the extension time of fingering observation is appropriately shortened.

[0225] From another perspective, high-scoring users can usually handle difficult techniques with faster finger movements, and their finger movements are precise, resulting in clearer movement characteristics. Appropriately shortening the extension time can effectively capture images of each individual triggered action.

[0226] The preset scoring threshold can be set by the user according to their practice needs. For example, a higher preset scoring threshold may indicate that the user has higher requirements for practice or that the performance is more difficult (such as multiple legato passages). In this case, the extended time can be set to be relatively longer, so as to more comprehensively (or for a longer period of time) identify the user's fingering movements.

[0227] In some embodiments, the steps further include:

[0228] The window update model is updated based on the performance index of the score, wherein the higher the performance index, the longer the first expansion time or the longer the second expansion time.

[0229] In some embodiments, performance metrics can be used to characterize the overall difficulty of a user’s performance of a musical score, providing a basis for adjusting the time window for the window update model.

[0230] In some embodiments, the performance indicators are defined and determined by at least one of performance rhythm, technical difficulty, and the user's performance ability.

[0231] For example, performance indicators can be ratings (such as 1-3 levels).

[0232] In some embodiments, performance metrics can be defined and determined directly by one of the following: performance rhythm, technical difficulty, or the user's performance ability. For example, the faster the performance rhythm, the higher the performance metric.

[0233] In some embodiments, the performance rhythm is characterized by the magnitude of an absolute velocity value, or the performance rhythm is characterized by the number of notes played per unit time.

[0234] In some embodiments, the difficulty of a technique can be determined based on the number of times or the time it takes for a typical user to practice the technique until it is mastered. For example, the more times a technique is practiced, or the longer the practice time, the higher its difficulty.

[0235] Alternatively, in some embodiments, the more techniques a score contains, the more difficult the techniques become. For example, a score requiring seven techniques is more difficult than a score requiring only one technique.

[0236] In some embodiments, the performance rhythm can be used to characterize the musical score's requirements for performance speed. The faster the performance rhythm, the higher the performance index, and the longer the corresponding extended time.

[0237] The unit of time can be the duration of a musical segment.

[0238] In some embodiments, the step of determining the performance index based on at least one of the performance rhythm, the technical difficulty, and the user's performance ability includes:

[0239] The performance difficulty level is determined based on the performance rhythm and / or the technical difficulty; wherein, the faster the performance rhythm, the higher the performance difficulty level; and the greater the technical difficulty, the higher the performance difficulty level.

[0240] Select a level threshold based on the performance ability;

[0241] Performance indicators are set according to the performance difficulty level and the level threshold; wherein, when the performance difficulty level is greater than the level threshold, the performance indicator is set as a first indicator, and when the performance difficulty level is less than or equal to the level threshold, the performance indicator is set as a second indicator; and the first indicator is higher than the second indicator.

[0242] In some embodiments, the user may choose to determine the difficulty level based on either the playing rhythm or the difficulty of the technique.

[0243] Alternatively, in some embodiments, a mapping table of different playing rhythms, different technical difficulties and playing difficulty levels can be preset in advance, and the corresponding playing difficulty level can be determined according to the mapping relationship.

[0244] It should be understood that the performance difficulty level reflects the actual difficulty of completing the score itself.

[0245] In some embodiments, a user's playing ability can be determined based on the user's performance score.

[0246] Alternatively, a user's playing ability can be obtained through user input.

[0247] In some embodiments, when a user's playing ability improves through a period of time (typically a longer period, such as a week), the level threshold can be updated accordingly (e.g., increased).

[0248] In some embodiments, a user's playing ability reflects the level of their playing technique. The stronger the playing ability, the higher the level threshold.

[0249] For example, the level threshold for beginners might be set at 3. A difficulty level above 3 would be considered quite difficult, or the performance index would be a high value, such as the first index. Conversely, the level threshold for experts might be set at 8. A difficulty level below 8 would be considered not very difficult, or the performance index would be a low value, such as the second index.

[0250] Correspondingly, for beginners, this invention can appropriately extend the time window, or combine more contextual information (such as redundant movements under extended time), to more comprehensively and for a longer period of time to achieve higher accuracy in recognizing their finger movements. For example, the first extended time of retrospection can show the subtle changes in joint angles, tendon tension, and even the preparatory sinking of the wrist as the finger moves from relaxation to pressing the key (which is crucial for determining which finger is exerting force), thereby effectively improving recognition accuracy. For advanced performers, the time window can be appropriately shortened to capture their rapid combinations, thus avoiding lower recognition accuracy due to a mismatch in the time window.

[0251] In some embodiments, while appropriately shortening the time window, the frame rate of the image within the corresponding time window can also be increased to ensure that each finger movement can be clearly captured.

[0252] Furthermore, by comparing the level threshold (user's subjective playing ability) and the playing difficulty level (objective difficulty of completing the score), this invention can better meet the real playing needs of different user groups. That is, when the playing index is high (such as when it is set as the first index), a longer time window is used to observe the user's fingering actions; when the playing index is low (such as when it is set as the second index), a shorter time window is used to observe the user's fingering actions, thereby making the allocation of computing resources more personalized.

[0253] In some embodiments, the present invention can also visualize the finger matching results. For example, the method further includes:

[0254] An image is displayed on a display device, and the finger matching results (such as for piano keys or fingers) are marked on the image.

[0255] The image may be an image captured by an image acquisition unit, or the image may be a virtual image obtained by fusing multiple viewpoint images.

[0256] The markings can include error markings. Error markings can be used to annotate the image of the finger and / or driving element in which the mistake was made during the performance. For example, they can be highlighted with prominent visual elements such as highlighted boxes, flashing icons, crosses, or specific lines.

[0257] In some embodiments, a pop-up window may also be displayed for specific error type descriptions (such as "fingering error", "key position deviation", "rhythm advance / delay").

[0258] The markings may include key markings. Key markings can be used to visually indicate at least one key involved in the current score and its triggering state. For example, specific keys that trigger or release events within the time window can be marked by highlighting, coloring, or outlining.

[0259] The markings can include finger markings. For example, visually indicating the fingers involved in triggering the driving elements on an image. For example, one or more fingers identified as triggering the keys can be precisely marked in an image using key point markings, outlines, or numbered labels (e.g., using numbers "1" to "5" to represent the thumb to the little finger, respectively) to clearly show the fingering relationships.

[0260] For example, referring to the patent application with publication number CN120526181A, markings can be added to information such as the triggered driving element and whether the fingering is correct.

[0261] From another perspective, this invention provides a motion capture method based on the synergy of sensor data and image data. This method enables precise digital analysis of a user's (e.g., a student's) playing performance using more accurate and reliable motion capture data. It provides users with more direct and accurate data feedback while significantly assisting piano teachers in completing professional teaching tasks. Furthermore, this motion capture method, which combines sensor and image data, reduces the technical difficulty of data processing while improving data reliability.

[0262] It is worth noting that, taking the auxiliary teaching scenario as an example, the motion capture method of this invention can analyze the user's playing state in real time (e.g., whether the fingering is accurate). However, real-time analysis places extremely high demands on computing power. To address this, this embodiment employs sensors for precise data acquisition, combined with image processing as a posture-assisted verification method. This approach improves the accuracy and reliability of motion capture data while reducing the performance requirements for real-time analysis to some extent through relatively less image processing.

[0263] Furthermore, by designing a collaborative recognition mode that prioritizes sensors and uses image processing as an auxiliary method, the real-time processing range of images can be expanded without excessively increasing the burden on image processing. For example, it can perform relatively complete recognition of multiple images with a wider temporal field of view (such as covering multiple moments before, during, and after the keys are pressed) in a short period of time, which also helps to improve the reliability of images as an auxiliary verification method. This avoids the difficulty in complete, clear, or reliable finger recognition caused by excessively fast hand speeds during complex playing techniques.

[0264] Alternatively, from another perspective, this embodiment preferably prioritizes the time result generated by the sensor (such as the timestamp of the moment of pressing), selecting the temporal field of view (or timeline) for image processing based on the sensor's time. This dual coordination of the time dimension and the functional task dimension can effectively balance the accuracy and computational burden of real-time analysis.

[0265] In some embodiments, see Figure 5 The present invention also provides an event recognition system applied to a musical instrument system, the musical instrument system comprising: a musical instrument, the musical instrument comprising: a driving element, which, when manipulated, can directly or indirectly control the musical instrument to produce sound; a sensor, the sensor being used to detect the manipulation state of the driving element, the manipulation state including: triggering and releasing; an image acquisition module, the image acquisition module comprising: at least two image acquisition units positioned facing the musical instrument, the at least two image acquisition units being used to acquire at least two images, and the at least two images having different viewing angles; correspondingly, the system comprises: a trigger judgment module, used to determine whether the driving element has triggered an event; when the determination result of the trigger judgment module is yes, then entering a time acquisition module: the time acquisition module is used to acquire the first moment when the driving element triggers the event, and the second moment when the driving element releases the event sequentially; a window setting module, used to set a corresponding time window for the trigger event using a window update model based on the first moment and the second moment; wherein, the window update model includes: the time window = [First moment - λ1, Second moment + λ2]; where λ1 is the set first extension time and λ2 is the set second extension time; Attention update module, used to increase or maintain the attention of the image acquisition module to the first attention level under the time window;

[0266] The system further includes a window update module; the window update module further includes: a performance type acquisition unit for acquiring the performance type; a recommended extension time selection unit for selecting at least one recommended extension time from the first extension time and the second extension time according to the performance type; and a window update unit for updating the window update model according to at least one of the recommended extension times; wherein the image acquisition module is used to generate finger matching results based on the at least two images, the finger matching results including: the matching relationship between at least one finger and the driving element.

[0267] It should be understood that the event recognition system applied to the musical instrument system can be used to implement the method steps described in any embodiment of the present invention.

[0268] In some embodiments, the present invention also provides a computer device, the device including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the method described in any embodiment of the present invention.

[0269] In some embodiments, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the method as described in any embodiment of the present invention.

[0270] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0271] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a computer terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0272] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A fingering recognition method applied to a musical instrument system, characterized in that, The musical instrument system includes: a musical instrument, the musical instrument including: a driving element, which, when manipulated, can directly or indirectly control the musical instrument to produce sound; a sensor, which is used to detect the manipulation state of the driving element, the manipulation state including: triggering and releasing; and an image acquisition module, the image acquisition module including: at least two image acquisition units arranged facing the musical instrument, the at least two image acquisition units being used to acquire at least two images, and the at least two images having different viewing angles; Correspondingly, the method includes the following steps: S101, determine whether a trigger event has occurred in the driving element; If S101 is true, then S102 is executed: S102, Collect the first moment when the triggering event occurs in the driving element, and the second moment when the release event occurs in sequence in the driving element; S103, a window update model is used to set a corresponding time window for the triggering event based on the first time point and the second time point; wherein, the window update model includes: The time window = [first moment - λ1, second moment + λ2]; where λ1 is the set first extended time and λ2 is the set second extended time. S104, within the time window, the attention of the image acquisition module is increased or maintained at the first level of attention; The image acquisition module is used to generate a finger matching result based on the at least two images. The finger matching result includes the matching relationship between at least one finger and the driving element.

2. The method according to claim 1, characterized in that, It also includes the following steps: S105, update the attention to a second attention level within an interval window; wherein the interval window is a time period during which the triggering event has not occurred; the second attention level is less than the first attention level.

3. The method according to claim 2, characterized in that, The interval window is the time period between two adjacent time windows.

4. The method according to claim 1, characterized in that, The musical instrument is a piano, and the driving element is the piano key.

5. The method according to claim 1, characterized in that, Before S104, the following steps are also included: When at least two adjacent trigger events are identified, the positions of the corresponding at least two driving elements and the corresponding at least two time windows are obtained; When the positional interval between at least two of the driving elements is less than a preset spacing threshold, the at least two time windows are merged into one time window.

6. The method according to claim 1, characterized in that, It also includes the following steps: When adjacent release events and trigger events are identified sequentially, the adjacent difference between the times corresponding to the release event and the trigger event is calculated; When the adjacent difference is greater than a preset time interval, the first extension time is set to the first extension value, and / or the second extension time is set to the second extension value; When the adjacent difference is less than the time interval, the first extended time is set to the third extended value, and / or the second extended time is set to the fourth extended value; Wherein, the first extended value is less than the third extended value, and the second extended value is less than the fourth extended value.

7. A fingering recognition system applied to a musical instrument system, characterized in that, The musical instrument system includes: a musical instrument, the musical instrument including: a driving element, which, when manipulated, can directly or indirectly control the musical instrument to produce sound; a sensor, which is used to detect the manipulation state of the driving element, the manipulation state including: triggering and releasing; and an image acquisition module, the image acquisition module including: at least two image acquisition units arranged facing the musical instrument, the at least two image acquisition units being used to acquire at least two images, and the at least two images having different viewing angles; Correspondingly, the system includes: A trigger judgment module is used to determine whether a trigger event has occurred in the driving element. When the judgment result of the trigger judgment module is yes, the time acquisition module is entered: the time acquisition module is used to acquire the first moment when the driving element triggers the event and the second moment when the driving element releases the event in sequence; a window setting module is used to set a corresponding time window for the trigger event using a window update model based on the first moment and the second moment; wherein, the window update model includes: the time window = [first moment - λ1, second moment + λ2]; wherein, λ1 is the set first extension time and λ2 is the set second extension time; an attention update module is used to increase or maintain the attention of the image acquisition module to a first attention level under the time window; wherein, the image acquisition module is used to generate a finger matching result based on the at least two images, and the finger matching result includes: the matching relationship between at least one finger and the driving element.

8. The system according to claim 7, characterized in that, The system also includes: An interval window module is used to update the attention to a second attention level within an interval window; wherein the interval window is a time period during which the triggering event has not occurred; and the second attention level is less than the first attention level.

9. A computer device, characterized in that, The device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, when executing the computer program, implement the fingering recognition method for a musical instrument system as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to implement a fingering recognition method for a musical instrument system as described in any one of claims 1 to 6.