A piano key recognition method, intelligent musical instrument, device and medium

Through staged interference filtering and perspective deformation correction, the robustness and accuracy issues of key recognition in complex scenes are solved, and high-precision key recognition effect is achieved.

CN120125810BActive Publication Date: 2025-09-19GRANMUS STAFF TECHNOLOGIES (CHONGQING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510585626.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-19
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Existing piano key recognition technology lacks robustness in complex scenarios and is affected by environmental noise such as natural light, stage lights and moving objects, resulting in reduced recognition accuracy, and changes in viewing angle increase the difficulty of recognition.

Method used

A phased and progressive interference filtering mechanism under strict time sequence is adopted. Through multi-directional smoothing, preset threshold processing and morphological operations combined with perspective deformation correction, image noise is gradually eliminated, key details are retained and recognition accuracy is improved.

Benefits of technology

Achieve high-precision and high-robust piano key recognition in complex scenarios, maintain key details and improve recognition accuracy and anti-interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125810B_ABST
    Figure CN120125810B_ABST
Patent Text Reader

Abstract

The present application relates to the field of digital performance, and in particular to a piano key recognition method, intelligent musical instrument, device, and medium. The method comprises: acquiring a first field of view image; performing a first filtering operation on the first field of view image to suppress a first form of image noise, thereby obtaining a second field of view image; performing a preset threshold processing on the second field of view image to obtain a third field of view image, wherein the third field of view image includes a plurality of target pixel sets, and the target pixel sets are composed of a plurality of continuous pixel points whose pixel values ​​are lower than a preset threshold; performing a second filtering operation on the target pixel sets in the third field of view image based on a preset structural element to suppress a second form of image noise, thereby obtaining a fifth field of view image; and identifying a plurality of first-class piano keys from the fifth field of view image based on a second preset model. In this way, high-precision interference elimination is achieved while retaining the details of the piano keys, effectively improving the recognition accuracy and robustness of the piano keys in complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of digital performance, and in particular to a piano key recognition method, intelligent musical instrument, device and medium. Background Art

[0002] With the development of science and technology, the digital expression of music and art has become increasingly diverse, with broad application prospects in education, entertainment, virtual reality, and other fields. Key recognition technology is a core technology in music teaching, performance assistance, and intelligent instrument development. Its core goal is to detect the position and status of piano keys through sensors or visual systems.

[0003] For example, an invention patent application with application publication number CN111695499A discloses a piano key recognition method, device, electronic device and storage medium, the method comprising: acquiring a first keyboard image of a target musical instrument, wherein the keyboard of the target musical instrument is provided with black keys and white keys; inputting the first keyboard image into a preset image recognition model to determine the outline of each black key in the first keyboard image; determining the key range corresponding to each octave in the first keyboard image and the outline of each white key in each octave based on the outline of the black keys and the spacing between adjacent black keys.

[0004] For example, the Chinese invention patent application with application publication number CN113723264A discloses a method for intelligently identifying playing errors to assist piano teaching, including: obtaining a 2D image of a piano being played containing a complete piano keyboard from above the piano keyboard; performing target detection on the 2D image through a piano keyboard detection network to detect the piano keyboard area represented by the relative position coordinates of the 2D image, and obtaining the piano keyboard position coordinates under the original coordinates of the 2D image through conversion.

[0005] However, in practical applications, natural light or stage lighting, moving objects and other environmental noises will affect the quality of the piano key images captured by the visual system, resulting in a decrease in the accuracy of piano key recognition and insufficient robustness in complex scenes. Summary of the Invention

[0006] The main purpose of this application is to provide a method for identifying piano keys, an intelligent musical instrument, a device, and a medium. In order to solve the above-mentioned technical problems, this application specifically adopts the following technical solutions:

[0007] A first aspect of the present application is to provide a method for identifying a piano key, the method comprising:

[0008] S301, acquiring a first field of view image, wherein the first field of view image includes images corresponding to a plurality of piano keys of a type;

[0009] S302, performing a first filtering operation on pixel points in the first field of view image based on pixel values ​​in the first field of view image to suppress image noise of a first form, thereby obtaining a second field of view image;

[0010] S303, performing preset threshold processing on the second field of view image to obtain a third field of view image, wherein the third field of view image includes a plurality of target pixel sets, each of which is composed of a plurality of consecutive pixel points whose pixel values ​​are lower than a preset threshold;

[0011] S304, performing a second filtering operation on the target pixel set in the third field of view image based on a preset structure element to suppress the second form of image noise, thereby obtaining a fifth field of view image;

[0012] S305 : Based on the second preset model, identify a plurality of keys of the type from the fifth field of view image.

[0013] A second aspect of the present application is to provide an intelligent musical instrument, comprising:

[0014] An image acquisition module, configured to acquire a first field of view image, wherein the first field of view image includes images corresponding to a plurality of piano keys of a type;

[0015] memory for storing computer programs;

[0016] A processor is used to execute the computer program and implement the steps of the key recognition method provided in any embodiment of the present application when executing the computer program.

[0017] A third aspect of the present application is to provide a computer device, comprising:

[0018] memory for storing computer programs;

[0019] A processor is used to execute the computer program and implement the steps of the key recognition method provided in any embodiment of the present application when executing the computer program.

[0020] The fourth aspect of the present application is to provide a corresponding computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor performs the steps of the key recognition method provided in any embodiment of the present application.

[0021] Beneficial effects:

[0022] This application proposes a piano key recognition method, intelligent musical instrument, device and medium, and specifically provides a phased and progressive interference filtering mechanism under strict timing. The filtering intensity of each stage is strictly limited, and only specific interference information is processed. In the early stage, the original image data is highly respected for fine-tuning, and in the later stage, adaptive correction is performed based on the reliable image data and camera parameters after processing. During the process, some interference information is tolerated to be transmitted backward and then gradually eliminated, avoiding the loss of details caused by one-time high-sensitivity filtering. In this way, high-precision interference removal is achieved while retaining the details of the piano keys, effectively improving the recognition accuracy and robustness of the piano keys in complex scenarios.

[0023] In the first stage, only discrete isolated noise points (such as sensor noise or dust reflections) are locally smoothed. The multi-directional smoothing strategy based on pixel distribution characteristics can keep the overall structure of the image unchanged and avoid excessive denoising leading to loss of details. The edge blurring caused by the smoothing strategy is clearly restored during threshold processing, ensuring the geometric integrity of the keys and providing a reliable parameter calibration basis for the second stage.

[0024] In the second stage, pixel aggregation is performed through conservative parameter settings to fill in the micro-fractures inside the black keys caused by the shooting angle (for example, micro-fractures inside the black keys caused by wood texture or uneven lighting), reconstruct the complete closed area, and then implement refined boundary trimming to simultaneously eliminate structural noise (such as burrs on the edges of the keys, linear interference caused by keyboard key gaps and reflections) and background interference, thereby achieving simultaneous optimization of repairing effective features and stripping invalid features, avoiding the generation of multiple broken edges in unclosed contours during the subsequent extraction process, which in turn leads to confusion in the contour hierarchy relationship.

[0025] In the third stage, the key contour is adaptively corrected based on perspective deformation. By processing the keys near and far from the optical axis in different regions and correcting the dynamic parameters, the calibrated key contour is ensured to restore the deformation error caused by the shooting angle, so as to accurately locate the keys at each position in the image. Even when the instrument model and shooting angle change, it can still maintain a high recognition rate and has stronger generalization ability.

[0026] The fourth stage performs hard physical verification of the key layout parameters. To avoid premature pruning that may cause the valid contour to be accidentally deleted due to local deformation, physical verification is performed at the end of the process. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the embodiments or the description of the prior art. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the various elements or parts are not necessarily drawn according to the actual scale. Obviously, the drawings described below are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without paying any creative work.

[0028] Figure 1 is a schematic flow chart of a piano key recognition method provided in an embodiment of the present application;

[0029] Figure 2 is a schematic diagram of a first field of view image provided in an embodiment of the present application;

[0030] Figure 3 is a schematic diagram of a second field of view image provided in an embodiment of the present application;

[0031] Figure 4 is a schematic diagram of a third field of view image provided in an embodiment of the present application;

[0032] Figure 5 is a schematic diagram of a fifth field of view image provided in an embodiment of the present application;

[0033] Figure 6 This is a schematic diagram of a type of piano key contour extraction provided by an embodiment of the present application;

[0034] Figure 7 This is a schematic diagram of another type of piano key wheel extraction provided in an embodiment of the present application;

[0035] Figure 8 This is a schematic diagram of a type of piano key contour after adjustment provided by an embodiment of the present application;

[0036] Figure 9 is a schematic flow chart of a method for identifying piano keys based on key seams provided in an embodiment of the present application;

[0037] Figure 10 is a schematic diagram of capturing a first partial image provided by an embodiment of the present application;

[0038] Figure 11 Schematic diagram of a method for identifying piano keys based on key seams provided in an embodiment of the present application;

[0039] Figure 12 is a schematic diagram of a piano key recognition result provided by an embodiment of the present application;

[0040] Figure 13This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] To make the purpose, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0042] Herein, suffixes such as "module," "component," or "unit" used to represent elements are only used to facilitate the description of the present application and have no specific meaning. Therefore, "module," "component," or "unit" can be used interchangeably.

[0043] As used herein, terms such as "upper," "lower," "inner," "outer," "front," "back," "one end," and "the other end" indicate positions or locations based on those shown in the accompanying drawings. These terms are intended solely to facilitate the description of this application and simplify the description. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0044] As used herein, unless otherwise expressly specified or limited, the terms "installed," "provided with," "connected," etc., should be understood broadly. For example, "connected" may refer to a fixed connection, a detachable connection, or an integral connection; it may refer to a mechanical connection, a direct connection, an indirect connection through an intermediate medium, or internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this application.

[0045] As used herein, "and / or" includes any and all combinations of one or more of the associated listed items.

[0046] Herein, "plurality" means two or more than two, ie, it includes two, three, four, five, etc.

[0047] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0048] In actual applications, the environment in which piano keys are located is highly variable, and image acquisition modules can be placed in a variety of locations. Current piano key recognition technology has low environmental adaptability, making it difficult to maintain high reliability in practical applications, limiting its widespread adoption. On the one hand, changes in natural light or stage lighting introduce additional noise, which can cause the images captured by the image acquisition module to be overexposed, reflective, or distorted, reducing feature extraction accuracy. Furthermore, complex environmental backgrounds (such as moving objects and cluttered backgrounds) can cause the target and background to become mixed, increasing recognition difficulty and significantly reducing recognition accuracy. Furthermore, the different placements of the image acquisition module can lead to changes in perspective, causing the piano keys to appear in different states in the image, further increasing recognition difficulty.

[0049] Based on this, the present application proposes a piano key recognition method, which specifically provides a phased and progressive interference filtering mechanism under strict timing. The filtering intensity of each stage is strictly limited, and only specific interference information is processed. In the early stage, the original image data is highly respected for fine-tuning, and in the later stage, adaptive correction is performed based on the processed reliable image data and camera parameters. During the process, some noise or blurred information is tolerated to be passed to subsequent stages, and various interference information in the image is gradually eliminated to avoid the loss of details caused by one-time high-sensitivity filtering. In this way, high-precision interference removal is achieved while retaining the details of the piano keys, effectively improving the recognition accuracy and robustness of the piano keys in complex scenarios.

[0050] Furthermore, this application proposes a key recognition method based on key gaps. This method uses a key with a distinct feature as a guide, quickly localizes the key gaps within the local range of the second key to be identified, and then normalizes the key gaps within the local range using deformation compensation using dynamic structural elements. This method then reversely infers the key's regional contours through the key gaps. Thus, through a multi-level convergence strategy of "global image-local image-key gaps," combined with a coupling mechanism of dynamic structural elements and perspective deformation parameters, this method achieves both efficient utilization of computing resources and optimization of recognition accuracy, effectively improving the recognition's anti-interference capability.

[0051] In this article, a musical instrument has keys arranged in a fixed order, and each key has a fixed pitch, and these keys can form a keyboard. The musical instrument can specifically be a piano, organ, accordion or electronic keyboard, etc., which is not limited here.

[0052] In this article, an image acquisition module is a device that converts optical signals in the physical world into digital images, such as a camera or camcorder. Its core components may include a photoelectric converter (such as a CCD or CMOS sensor), a lens, a signal processing unit, and a storage module. The model and specific components of the image acquisition module can be selected according to actual needs and are not limited here.

[0053] For example, the installation location of the image acquisition module can be determined based on the specific application scenario (such as ambient light) and technical parameters (such as resolution, installation height, angle, image integrity, etc.) to ensure device performance and compliance. The specific installation location is not limited here. For example, the image acquisition module can be perpendicular to the plane of the piano keys and located directly above the keyboard, such as on the inside of the piano lid, the top of the piano frame, or a hanging bracket. For another example, the image acquisition module can be installed at a certain angle to the plane of the piano keys and on the front or upper side of the piano. For another example, the image acquisition module can be independently set at a position corresponding to the user's perspective.

[0054] In this paper, perspective is a fundamental law in optical imaging, describing the deformation of three-dimensional objects in two-dimensional images, where they appear larger when closer and smaller when farther away due to distance differences. Correspondingly, the perspective deformation parameters reflect the deformation characteristics of the piano keys caused by the camera's viewing angle, such as the shrinkage ratio and direction. Positions close to the image acquisition module have a lower degree of perspective than those at the edge, and the corresponding perspective deformation parameters also change accordingly.

[0055] It should be understood that the image acquisition module has preset internal and external parameters, namely, preset camera parameters, including focal length, mounting height, pitch and roll angles, sensor size, lens distortion coefficient, etc., which can be used to calibrate the optical axis position. The object's perspective deformation parameters are determined based on these preset camera parameters and the object's position in the image (e.g., the coordinates of the position of a first-class piano key or a second-class piano key). For example, the object's perspective deformation parameters are determined based on the angle parameter (i.e., the perspective angle) between the object's position in the image and the optical axis of the image acquisition module. Positions closer to the optical axis have a smaller perspective angle than those at the edge, resulting in a correspondingly lower degree of deformation.

[0056] In this article, morphological operations are a set of mathematical methods for processing digital images, primarily used to extract specific shape information from image components. Based on set theory principles, morphological operations include dilation, erosion, opening, and closing, used to achieve functions such as noise removal and enhancing or weakening shape features. Structuring elements are a core tool in mathematical morphology for image processing. Essentially, they are a set of pixels with known shape, size, and orientation, defining the shape and scope of the morphological operations.

[0057] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features of the embodiments can be combined with each other. Figure 1 , Figure 1 is a schematic flow chart of a method for identifying a piano key provided in an embodiment of the present application, such as Figure 1 As shown, an embodiment of the present application provides a method for identifying piano keys; the method includes S301 to S305.

[0058] S301 : Acquire a first field of view image, where the first field of view image includes images corresponding to a plurality of piano keys of a type.

[0059] The first field of view image refers to the original image directly captured by the image acquisition module. The image needs to contain a type of piano keys that need to be identified, serving as the input basis for subsequent key recognition.

[0060] See also Figure 2 , Figure 2 is a schematic diagram of a first field of view image provided in an embodiment of the present application, such as Figure 2 As shown, the musical instrument is a piano, and the image acquisition module is vertically arranged on the plane of the piano keys and located directly above the keyboard. The first field of view image includes several first-class piano keys 100 to be identified (such as the black keys in the figure), and may also include other parts of the instrument, such as the casing, strings, hammers, etc., as well as the background of the venue where the instrument is located, the hand movements of the performer, etc. It can be seen that in the actual performance scene, there are complex interfering environmental factors, such as Figure 2 The reflection of the piano keys, the messy background, etc.

[0061] S302 : Perform a first filtering operation on pixel points in the first field of view image based on pixel values ​​in the first field of view image to suppress image noise of a first form, thereby obtaining a second field of view image.

[0062] Specifically, the first form of image noise refers to discrete, isolated interference within an image, such as random, discrete noise points generated by the sensor, dust reflections, or tiny light spots in the environment. This type of noise appears as interfering points in the image that differ significantly from surrounding pixel values ​​but are spatially dispersed. This can disrupt the integrity and continuity of key structures, such as the outline of piano keys.

[0063] Correspondingly, the first filtering operation is an image processing method that performs adaptive smoothing on a local area based on the grayscale or color distribution characteristics of each pixel point in the first field of view image. For example, by analyzing the pixel value distribution within the pixel neighborhood, a multi-directional smoothing strategy is adopted to process the first field of view image, thereby suppressing the first form of image noise.

[0064] In some embodiments, only the first field of view image is selectively smoothed. For example, pixels suspected of being noise points (such as pixels whose difference from the neighborhood median exceeds a preset median threshold) are replaced with neighborhood statistical values ​​(such as median or weighted mean), while the original values ​​of the key edges are retained.

[0065] In some embodiments, the multi-directional smoothing strategy can be a median filter, a Gaussian filter or an adaptive weight filter. Preferably, the multi-directional smoothing strategy is a Gaussian filter. It should be understood that the Gaussian filter uses a Gaussian function to assign weights. Pixels close to the center are assigned higher weights, while pixels far from the center are assigned lower weights. This weight distribution mechanism makes the Gaussian filter more natural and smooth when removing noise, but it also cannot completely eliminate noise. Thus, the denoising effect is sacrificed to ensure that image details and key edges are retained.

[0066] It should be understood that the first stage only performs local smoothing on discrete, isolated noise points (such as sensor noise or dust reflections). The multi-directional smoothing strategy based on pixel distribution characteristics can maintain the overall structure of the image, avoid excessive denoising and loss of detail, ensure the geometric integrity of the keys, and thus provide a reliable basis for parameter calibration in the second stage. Because the interference filtering strength in the first stage is strictly limited, the second field of view image generated after the first stage processing only has a preliminary cleansing of discrete noise. While it retains the clarity of the main structure of the keys, it also retains complex environmental interference such as local fractures caused by shooting angle, uneven lighting, or black key texture. Please refer to Figure 3 , Figure 3 is a schematic diagram of a second field of view image provided in an embodiment of the present application, such as Figure 3 As shown in the figure, from a macroscopic perspective, there are still environmental interferences such as piano key reflections and background clutter in the second field of view image.

[0067] In some embodiments, S302 includes: identifying at least one pixel curve on the first field of view image, the pixel curve consisting of multiple pixel points, and the pixel curve is used to reflect the degree of change of the pixel value of the pixel point; calculating the pixel difference degree between at least one first target segment and an adjacent second target segment in the pixel curve, wherein the first target segment and the second target segment are composed of at least one pixel point; determining whether the pixel difference degree exceeds a pixel difference threshold; if so, correcting the first target segment to reduce the pixel difference degree.

[0068] Specifically, at least one pixel curve in the first field of view image is identified. The pixel curve is composed of continuously arranged pixel points, and its pixel values ​​show dynamic changes along the curve path, such as light and dark transitions or texture trends. Furthermore, the pixel curve can be segmented with a fixed length or adaptively to obtain target segments of several continuous pixel points. For any first target segment, at least one second target segment adjacent to it is obtained, and the pixel difference degree of the target segment is calculated, such as mean difference, variance ratio or maximum gradient difference. If the pixel difference degree exceeds the pixel difference threshold, a noise area showing abnormal peaks or valleys as mutations is identified, and the first target segment is corrected, for example, by replacing it with a smooth transition value of the second target segment or interpolation reconstruction to eliminate the mutation and reduce the degree of difference. Among them, the pixel difference threshold is a preset difference tolerance critical value used to trigger the correction operation. The specific value can be determined based on noise feature statistics or empirical values ​​and is not limited here.

[0069] In some embodiments, if not, the original pixel value of the first target segment is retained to avoid redundant correction of normal pixels, thereby maintaining the integrity of the key edges and details.

[0070] In some embodiments, a plurality of pixel curves are generated along a specific direction in the first field of view image, which may be a horizontal direction, a vertical direction, or a tangent direction of a key edge, etc., which is not limited here.

[0071] Therefore, by analyzing the pixel value differences between adjacent areas segment by segment, we can distinguish target segments that present continuous or step changes, accurately locate and suppress the first morphological noise, avoid over-correction that leads to blurred key contours, and provide a clearer intermediate image for subsequent threshold processing.

[0072] S303 , performing preset threshold processing on the second field of view image to obtain a third field of view image, wherein the third field of view image includes a plurality of target pixel sets, and the target pixel sets are composed of a plurality of continuous pixel points whose pixel values ​​are lower than a preset threshold.

[0073] The preset threshold processing is a processing step that converts the grayscale or color information of the second field of view image into a binary image. The image pixels are divided into foreground and background by setting a fixed threshold or a dynamically calculated adaptive threshold. Preferably, the preset threshold processing is an adaptive binarization corresponding to the adaptive threshold. The threshold is dynamically calculated based on the regional illumination to adapt to uneven illumination or reflective interference. An illumination normalization layer is added before the adaptive binarization. Contrast Limited Adaptive Histogram Equalization (CLAHE) is used to enhance local contrast without amplifying noise, thereby achieving adaptive illumination compensation.

[0074] Specifically, the second field of view image is processed by a preset threshold, and the obtained third field of view image is a binary image. The pixels below the preset threshold in the third field of view image are marked as target pixels, and the adjacent target pixels form a target pixel set. Figure 4 , Figure 4 is a schematic diagram of a third field of view image provided in an embodiment of the present application, such as Figure 4 As shown, taking black and white piano keys as an example, all continuous and adjacent "black" pixels (i.e., areas below a preset threshold) in the third field of view image form a target pixel set. Each black piano key will correspond to a target pixel set, allowing subsequent steps to more accurately identify the target pixel set as an independent target set while eliminating small interference.

[0075] It should be understood that using threshold segmentation further clarifies the previously preserved key outlines. For example, it can sharpen the slight edge blur caused by smoothing in step S302, suppress local grayscale interference caused by uneven lighting, and initially close microscopic fractures within the black keys caused by texture or uneven lighting. At the same time, it enhances the contrast between the keys and their surroundings, such as further differentiating the black and white keys on a piano, providing highly confident binary input for subsequent morphological operations.

[0076] In some embodiments, before step S303, the process further includes: increasing the chroma-weighted grayscale of the second field of view image, i.e., increasing the difference weights of the green and blue channels. It should be understood that, given the black-white color trend, at the interface between the black and white keys, the average color difference between green and blue is significantly higher than that between red and green, and red and blue. Furthermore, under strong light, the chroma difference between the black and white keys may be more significant than the brightness difference. Therefore, increasing the green and blue channels can maximize the color difference, thereby retaining more effective features in subsequent processing.

[0077] S304 : Perform a second filtering operation on the target pixel set in the third field of view image based on a preset structure element to suppress the second form of image noise, thereby obtaining a fifth field of view image.

[0078] The preset structuring element is a pre-set set of pixels used to detect and correct shape defects in the target pixel set. Its size and shape must match the physical characteristics of a class of piano keys, such as their width. Correspondingly, the second filtering operation is a morphological operation performed on the preset structuring element. The main operations include dilation, erosion, opening, and closing, used to achieve functions such as noise removal and strengthening or weakening shape features.

[0079] Specifically, the structured processing based on morphological operations uses preset structuring elements to repair the second form of image noise in the third field of view image. Second form image noise refers to non-discrete, continuous interference in the image caused by shooting conditions or physical structure. This includes small gaps within piano keys due to wood grain or uneven lighting, linear interference caused by key reflections, key seams between non-class keys, or dust accumulation.

[0080] It should be understood that in the second stage, pixel aggregation is used to fill the microscopic fractures inside the black key caused by the shooting angle, reconstruct the complete closed area, and then implement refined boundary trimming to simultaneously eliminate structural noise and background interference, thereby achieving simultaneous optimization of repairing effective features and stripping invalid features, avoiding the generation of multiple broken edges in unclosed contours during the subsequent extraction process, which in turn leads to confusion in the contour hierarchy relationship.

[0081] See also Figure 5 , Figure 5 is a schematic diagram of a fifth field of view image provided by an embodiment of the present application, such as Figure 5 As shown in the figure, after the second filtering operation, the fifth field of view image is processed, and the burrs on the edges of the keys (such as the bottom edge of the black keys), the key seams of the white keys, and multiple linear interferences in the background are eliminated, thereby achieving refined boundary trimming. At the same time, the interior of the black keys appears to be a complete, closed, and continuous state, as shown in the figure. Figure 4 As shown in the figure, the black key on the far right is interrupted by the reflective area, resulting in multiple breaks in the black key after the preset threshold processing, as shown in the figure. Figure 5 As shown, the gap caused by reflection on the rightmost black key in the fifth field of view image has been corrected to a complete, closed, continuous state, thereby avoiding the generation of multiple broken edges in the unclosed contour during the subsequent extraction process, which in turn leads to confusion in the contour hierarchy relationship. At the same time, the original geometric features of the keys are retained, that is, the rightmost black key is still in an incomplete form.

[0082] It should be understood that in order to ensure that the key details are fully preserved, the interference filtering strength of the second stage is still strictly limited, e.g. Figure 5 Many black keys are incomplete and the background still contains noise. This is because the second-stage interference filtering uses conservative parameter settings to avoid excessive interference and damage to the original geometric features of the keys, providing highly complete and authentic outline data for subsequent perspective correction and physical verification. Furthermore, both the first and second stages are preliminary processing, fine-tuning the original image data while highly respecting the original image data to eliminate specific image noise.

[0083] In some embodiments, the second form of image noise includes burr noise and linear noise; the S304 includes: based on the first structural element, expanding the boundaries of each target pixel set in the third field of view image to fill the internal holes or breaks of the target pixel set to obtain the fourth field of view image; based on the second structural element, shrinking the boundaries of each target pixel set in the fourth field of view image to eliminate the burr noise of each target pixel set and the linear noise in the fourth field of view image to obtain the fifth field of view image.

[0084] Specifically, the first structuring element is used to perform a boundary extension operation on the target pixel set in the third field of view image. Through morphological dilation, the gaps between adjacent pixels are filled to repair holes or breaks in the target pixel set caused by threshold processing or uneven lighting, such as local missing areas in the black key texture, thereby obtaining a fourth field of view image, in which the continuity of the target pixel set is enhanced, but the edges may be slightly redundant due to expansion. Furthermore, the second structuring element is used to perform a boundary contraction operation on the fourth field of view image. Through morphological erosion, the burr noise at the edge of the target pixel set and the residual linear noise in the image are eliminated. In the final generated fifth field of view image, the outline of the target pixel set is smoother and more compact, while retaining the overall structural characteristics of the piano keys. Among them, burr noise refers to small, irregular protrusions caused by preset threshold processing or uneven lighting, such as jagged defects on the edge of the piano keys. Linear noise refers to the slender striped noise in the image caused by light interference or sensor error, as well as other line interference in the background.

[0085] It should be understood that the preset structuring elements include a first structuring element and a second structuring element. To ensure that key details are fully preserved, more conservative first and second structuring elements are selected. For example, a specific design is performed based on the physical properties and noise characteristics of a particular type of key, with strict restrictions on the shape and size of its pixel set. Furthermore, the first and second structuring elements can be configured with different structures based on their different functions. The specific values ​​can be flexibly set based on actual conditions and are not limited here.

[0086] S305 : Based on the second preset model, identify a plurality of keys of the type from the fifth field of view image.

[0087] Specifically, by inputting a cleaned set of target pixels (e.g., a continuous, unbroken piano key region) from the fifth field of view image, the second preset model, combined with a pretrained feature library (e.g., key shape, size, and arrangement pattern) or a rule-based algorithm (e.g., contour matching, proportional analysis), accurately locates and classifies a class of keys that meet the predefined characteristics. For example, if the target is a black key, the model will identify narrow, long, and regularly arranged dark regions in the image. It should be understood that because the previous processing has significantly cleaned the image, the second preset model can identify a class of keys from the continuous set of target pixels using only simple rules or pretrained feature matching. Therefore, the second preset model is a preconfigured recognition algorithm or a trained classification model, whose feature library or rules are constructed based on prior information such as the geometry and arrangement of a class of keys. Alternatively, existing algorithms or contour matching models may be used, without limitation. The recognition result for a class of keys may include the key outline and / or the key's position coordinates in the image.

[0088] In some embodiments, since some noise or interference information is allowed to be passed to the subsequent stages in the aforementioned processing steps, the contour of a type of piano key extracted may have local defects, uneven contours, and other problems. Figure 6 , Figure 6 This is a schematic diagram of a type of piano key contour extraction provided by an embodiment of the present application, such as Figure 6 As shown, the red lines represent the extracted contours of a class of piano keys. Due to the significant effectiveness of the previous denoising stage, the second preset model accurately and completely extracts every class of keys (such as the black keys in the image). However, the contours of many of the black keys are not smooth; some keys (such as the rightmost black key) are partially misidentified. Based on this, this embodiment of the application proposes a contour correction mechanism based on perspective deformation.

[0089] In some embodiments, the method further includes: based on preset camera parameters, determining the perspective deformation parameters of each of the first-class keys according to the position coordinates of the first-class keys in the fifth field of view image; determining the target morphological parameters of each of the first-class keys according to the baseline morphological parameters of the first-class keys and the perspective deformation parameters, and adjusting the outline of the corresponding first-class keys according to the target morphological parameters of each of the first-class keys.

[0090] Specifically, based on preset camera parameters and the position coordinates of a class of keys in the fifth field of view image, the system determines the perspective deformation parameters (e.g., shrinkage ratio) for each key. Based on these perspective deformation parameters, the key's baseline morphological parameters are adjusted to simulate a perspective effect, causing the baseline morphological parameters to deform due to perspective, such as stretching or shortening the edges. Finally, the key contours are corrected based on the target morphological parameters of each key, ensuring that the calibrated key contours restore deformation errors caused by the shooting angle, thereby improving the accuracy of subsequent positioning.

[0091] It should be noted that for the same preset camera parameters, the perspective deformation parameters at each position coordinate are fixed. Based on the preset camera parameters and the position coordinates of a class of piano keys, the corresponding perspective deformation parameters can be quickly obtained. The perspective deformation parameters at each position coordinate in the image are determined based on prior data and are not limited here. Furthermore, the target morphological parameters for each piano key can be derived through affine or projective transformations.

[0092] In some embodiments, the reference morphological parameters may be geometric characteristic parameters of the piano key in the image without perspective distortion, such as width, length, aspect ratio, etc. The reference morphological parameters may be acquired in advance, or the morphological parameters of the piano key with the lowest degree of perspective distortion may be determined from the image and used as the reference morphological parameters.

[0093] In some embodiments, the method also includes: obtaining at least one proximal Class I key whose perspective deformation parameters are lower than a first deformation threshold, and determining the baseline morphological parameters of the Class I key based on the morphological parameters of at least one proximal Class I key; obtaining at least one distal Class I key whose perspective deformation parameters are higher than a second deformation threshold, wherein the first deformation threshold is less than or equal to the second deformation threshold; adjusting the baseline standard morphological parameters of the Class I key based on the perspective deformation parameters of each of the distal Class I keys, and determining the target morphological parameters of each of the distal Class I keys; and adjusting the contour of the corresponding distal Class I key according to the target morphological parameters of each of the distal Class I keys.

[0094] Specifically, a class of proximal keys (such as keys directly in front of the camera with almost no perspective distortion) with perspective deformation parameters lower than a first deformation threshold are screened out, and their shapes are close to the baseline shape under the standard viewing angle. By counting the morphological parameters of at least one proximal key, the baseline morphological parameters of a class of keys are calculated and updated as a reference for subsequent adjustments. For a class of distal keys (such as keys with obvious deformation due to tilt at the edge of the picture) with perspective deformation parameters higher than a second deformation threshold, the baseline morphological parameters are dynamically adjusted through perspective transformation based on their specific perspective deformation parameters. For example, if the width of a key is shortened due to tilting at a long distance, the baseline morphological parameters are adjusted to synchronously shorten the width to generate the target morphological parameters. Among them, the first deformation threshold is used to identify a class of proximal keys that are less affected by perspective deformation. Correspondingly, the second deformation threshold is used to identify distal keys affected by perspective deformation. The specific values ​​can be set flexibly and are not limited here.

[0095] Based on the target morphological parameters of each far-end key, its contour shape in the image is adjusted (such as stretching or scaling the pixel area in a specific direction) to make it closer to the perspective shape. Therefore, through near-end and far-end region processing and dynamic parameter correction, the key deformation caused by the difference in camera perspective is adapted during contour calibration, so as to accurately locate the keys at various positions in the image.

[0096] In some embodiments, when perspective does not occur, the scaling ratio in the perspective deformation parameter is that the target morphological parameter is 1.0 times the baseline morphological parameter. As a type of piano key gradually moves away from the optical axis in the image, the scaling ratio changes linearly, gradually reaching a target morphological parameter of 0.5 times the baseline morphological parameter.

[0097] In some embodiments, the outlines of several keys of a class are obtained, and the keys of the class are screened based on their widths. Keys of a class with widths greater than a threshold are obtained, and the keys of the class with the narrowest width are screened out. The morphological parameters of the keys of the class with the narrowest width are used as the reference morphological parameters. The width threshold can be flexibly set and is not limited here.

[0098] In some embodiments, the target morphology parameter is set with a limit value. When the target morphology parameter obtained by adjusting the baseline morphology parameter is less than the limit value, the limit value is used as the target morphology parameter. And / or, the number of abnormal times when the target morphology parameter obtained by adjusting the baseline morphology parameter is less than the limit value is counted. If the abnormal number is greater than a preset abnormal number, a proximal key of the first category is reselected to update the baseline morphology parameter. For example, if key A is selected as the proximal key of the first category, and the target morphology parameter obtained by adjusting the baseline morphology parameter of key A is less than the limit value, the limit value is used as the target morphology parameter. If the target morphology parameter is less than the limit value multiple times continuously, key B is reselected as the proximal key of the first category.

[0099] It should be understood that in the third stage, the morphological data of the proximal keys are combined as the reference parameters, and the distal keys are processed hierarchically through perspective deformation parameters, which not only maintains the reliability of parameter calibration, but also avoids parameter distortion caused by perspective projection. Even when there are differences in the production process of musical instruments, changes in instrument models or changes in shooting angles, a high recognition rate can still be maintained, and it has stronger generalization ability.

[0100] In some embodiments, see Figure 7 , Figure 7 This is a schematic diagram of another type of piano key wheel extraction provided by an embodiment of the present application, such as Figure 7As shown, the red lines represent the extracted outlines of the first-class piano keys. The second preset model accurately and completely extracts every first-class key. However, due to excessive background interference, the second preset model mistakenly identifies non-key areas as keys, such as the striped carpet and the long rectangular frame in the image. This misidentification can be further filtered out through position verification or arrangement pattern analysis in the fourth stage to ensure the accuracy of the final result. Therefore, the extraction of the first-class keys in step S305 also serves as a filter for image interference. The filtering strength at this stage is still strictly limited. In this case, high-quality fifth-field images are used for efficient initial screening, rather than pursuing absolute accuracy, and a certain degree of recognition error can be tolerated.

[0101] In some embodiments, the method further includes: obtaining the actual morphological parameters and position coordinate parameters of each Class I key; based on preset key layout parameters, screening several Class I keys according to the actual morphological parameters and position coordinate parameters of the Class I keys to determine several target Class I keys; and outputting the position coordinates of several target Class I keys in the first field of view image.

[0102] The preset key layout parameters are a set of standardized physical parameters for a class of keys in an image, used to verify the geometric and arrangement rationality of the recognition results. Exemplarily, the preset key layout parameters include at least one of a target length, target width, target aspect ratio, target area, target key spacing, and target number of keys for the class of keys. For example, the total number of black keys on a standard piano is 36.

[0103] Specifically, for each identified key, or each key after adaptively correcting its key outline based on perspective deformation, the actual morphological parameters and positional coordinate parameters are obtained. A multi-level verification process is then performed based on preset key layout parameters. First, individual key morphological verification is performed using the preset key layout parameters (target length, target width, target aspect ratio, target area, etc.). Areas that conform to the key's physical characteristics are strictly screened, and keys with significantly deviated size, aspect ratio, or area are eliminated, eliminating any abnormally sized interference objects. Second, the overall keyboard layout is verified based on key spacing, number, and arrangement. Keys that conform to the preset arrangement pattern and number of keys are retained, ensuring that the recognition results conform to the standard musical instrument structure. This ensures that target keys that meet the standard are accurately retained, while misidentified non-key areas in the background are eliminated. The output outputs the precise positional coordinates of the target keys in the original first-field-of-view image, providing reliable data for subsequent positioning or interactive operations.

[0104] In the embodiment of the present application, the interference filtering intensity of step S302 is strictly limited, and only the discrete noise is preliminarily purified. While retaining the clarity of the main structure of the keys, it also retains complex environmental interference such as local breaks caused by shooting angles, uneven lighting, or black key textures. It is strictly limited to complete the removal of discrete isolated noise points before threshold processing (step S303), ensuring that the edge blur caused by the smoothing strategy is clearly restored during threshold processing, ensuring the geometric integrity of the keys, and thus providing high-confidence binary input for subsequent morphological operations. Furthermore, step S304, while removing continuous noise, repairs the keys or edge breaks that appear after threshold processing, ensuring the closure of the key contour, and avoiding the generation of multiple broken edges in the unclosed contour during the subsequent extraction process, which leads to confusion in the contour hierarchy. Since the second stage of interference filtering uses conservative parameter settings, the processed keys may be incomplete and the background may still contain interference.

[0105] In step S305, the extraction of a class of piano keys is performed using high-quality fifth field of view images to achieve efficient initial screening, which can filter out most background interference. However, the filtering intensity is still strictly limited at this time, and absolute accuracy is not pursued. Instead, a certain degree of recognition error is tolerated, such as outline defects of a class of piano keys and misjudgment of non-key areas. Figure 8 , Figure 8 This is a schematic diagram of a type of piano key contour adjustment provided by an embodiment of the present application, such as Figure 8 As shown in Figure 2, these recognition errors are eliminated by adaptively correcting the key contour based on perspective deformation and performing hard physical verification based on key layout parameters. Figure 8 The green line in the middle represents the adjusted outline of a key. While accurately fitting a key, the outline maintains a smooth, standardized form. Furthermore, to avoid premature pruning that could result in the accidental deletion of valid outlines due to local deformation, physical verification is performed at the end of the process.

[0106] Therefore, the various steps in the embodiment of the present application are highly sequential and coupled. Through a phased and progressive interference filtering mechanism, even if some noise or fuzzy information is allowed to be passed to subsequent stages, it will be gradually identified and corrected in subsequent steps. Ultimately, high-precision interference removal is achieved while retaining the details of the keys, effectively improving the recognition accuracy and robustness of the keys in complex scenarios.

[0107] See also Figure 9 , Figure 9 is a schematic flow chart of a method for identifying piano keys based on key seams provided in an embodiment of the present application, such as Figure 9 As shown, an embodiment of the present application provides a method for identifying piano keys based on key gaps; the method includes S401 to S404.

[0108] S401 : Acquire a first field of view image, where the first field of view image includes images corresponding to a plurality of first-category piano keys and second-category piano keys.

[0109] Among them, the first field of view image refers to the original image directly captured by the image acquisition module. The image presents both type I and type II keys and must contain sufficient details to support subsequent positioning and segmentation operations, such as clear key seam boundaries or key contours, as the input basis for subsequent key recognition.

[0110] S402 , based on the preset positional relationship between the first type of keys and the second type of keys, and according to the key image area where the first type of keys are located, intercepting first partial images of the area where a plurality of second type of keys are located from the first field of view image.

[0111] Among them, the preset position relationship refers to the fixed spatial arrangement rules between the first type of keys and the second type of keys. Taking the piano as an example, the black keys are usually located in the gap above the adjacent white keys, and there is a fixed size ratio and specific spacing arrangement between the two, such as two black keys sandwiched between three white keys.

[0112] Specifically, based on the known key image region where the first type of keys reside, a first partial image containing only the second type of keys is extracted from the first field of view image according to a preset positional relationship. This rapidly narrows the processing scope through spatial correlation, converging global processing to a localized area, reducing computational effort and eliminating irrelevant background interference, while preserving the complete details of the target area and providing high-precision input for subsequent key gap recognition.

[0113] See also Figure 10 , Figure 10 is a schematic diagram of capturing a first partial image provided by an embodiment of the present application, such as Figure 10 As shown, the musical instrument is a piano, and the image acquisition module is positioned perpendicular to the plane of the piano keys and directly above the keyboard. The first field of view image includes several first-class piano keys 100 (black keys in the image) and several second-class piano keys 200 (white keys in the image). Based on the preset positional relationship, rectangular areas are cut below and on both sides of the first-class piano keys 100, as shown by the red dashed lines, to cover the area where the second-class piano keys 200 may be located.

[0114] See also Figure 11 , Figure 11 Schematic diagram of a method for identifying piano keys based on key seams provided in an embodiment of the present application, such as Figure 11As shown in 40a, a region of interest (ROI) of the white keys, i.e., a first partial image, is captured. This includes partial images of several second-class keys, quickly eliminating irrelevant background interference. It should be understood that there are significant differences in color, shape, or texture between the first-class keys and the second-class keys, such as a contrast between dark and light colors. Therefore, the first partial image can be quickly captured based on the first-class keys.

[0115] In some embodiments, shape features (such as rectangles and aspect ratios) can be combined to filter regions that match the size and shape of a key of the first type. Alternatively, template images of the first type of key (such as local textures or outlines) can be pre-stored and then slidingly matched within the target image to find the region with the highest similarity. A detector for the first type of key can also be trained using a large amount of annotated data to output its location and category. The features of the first type of key are more distinct than those of the second type, resulting in higher recognition accuracy. Methods for identifying the first type of key can be found in related art and are not limited here.

[0116] In some embodiments, a key recognition method provided in any embodiment of the present application can be used to identify a type of keys and the position coordinates of the type of keys in the first field of view image, and then determine the key image area where the type of keys are located.

[0117] In some embodiments, the keys on a common musical instrument keyboard are divided into two categories: black keys and white keys. The black keys are considered as category one keys, and the white keys are considered as category two keys. Preferably, a first partial image is captured using the black key coordinates as a reference, and white key recognition is performed based on the first partial image. Compared to white keys, the features of black keys are more distinct, so global processing can avoid introducing a large amount of noise and improve computational efficiency. Furthermore, black key projections and white key edges are difficult to distinguish in the global image, so local processing of white keys can reduce the false detection rate.

[0118] In some embodiments, the preset positional relationship is determined based on prior knowledge of standard musical instruments. Exemplarily, the preset positional relationship between the first type of keys and the second type of keys is a shape ratio or area ratio between the first type of keys and the second type of keys. The shape ratio is the relative geometric relationship between the two types of keys, such as the ratio of length, width, or height. For example, the length of a white key on a piano is generally two-thirds the length of a black key. The area ratio is the ratio of the areas of the two types of keys in an image or in their actual physical structure.

[0119] In some embodiments, after S402, it also includes: performing preset threshold processing on the first local image, and performing a first filtering operation to suppress the first form of image noise; or, performing a first filtering operation on the first local image to suppress the first form of image noise, and performing preset threshold processing.

[0120] Specifically, the preset threshold processing is a binarization processing corresponding to a global or adaptive threshold, which separates the piano keys in the first local image from other backgrounds (such as key seams); the first filtering operation is to suppress the image noise of the first form through a multi-directional smoothing strategy. Preferably, the multi-directional smoothing strategy is Gaussian filtering. The specific implementation method can refer to the aforementioned embodiment and will not be repeated here. The specific parameter settings of the preset threshold processing and the first filtering operation can be the same as or different from the aforementioned embodiment and are not limited here. In addition, since the background interference in the first local image is extremely low, and the key seam features are clear and not easy to blur, the two processing can be flexibly combined in sequence and are not limited here. The first local image obtained by processing is as follows. Figure 11 As shown in 40b.

[0121] S403 , identifying a plurality of initial inter-key gaps between adjacent two types of keys in the first partial image; and performing a morphological operation on the plurality of initial inter-key gaps in the first partial image to obtain a target partial image.

[0122] Specifically, the initial inter-key gaps between adjacent two types of keys are identified from the first partial image, such as low-brightness areas or boundary lines. Figure 11 40c is the initial key gap recognition result, where the lines corresponding to the white pixels are the initial key gaps. Further, morphological operations are performed on these initial key gaps to eliminate structural noise and standardize the gap morphology. Figure 11 40d is the target local image output after morphological operation processing. The gaps between the keys have been corrected to continuous strips that conform to physical laws (such as uniform width and vertical direction). At the same time, the arrangement relationship between the key gaps and the keys is retained, providing high-confidence input for the key area extraction in S404.

[0123] It should be understood that the gap between keys (also called key gap) is the natural boundary area between two adjacent types of keys in the first partial image, which appears as a narrow strip with low brightness or high contrast. Its recognition method includes local feature analysis based on color contrast (such as the brightness difference between the key gap and the key) or edge detection. Figure 11 As shown in FIG40c, the image region corresponding to the initial gap between keys is very narrow, and the morphological processing thereof needs to strictly limit the corresponding structural elements to avoid the gap being overfilled or destroying the original gap information.

[0124] In some embodiments, the S403 includes: identifying the contours of several of the second-category keys from the first partial image based on the color information of several pixels in the first partial image; identifying several candidate inter-key gaps between the two-category keys from the first partial image; obtaining at least one adjacent second-category key adjacent to the candidate inter-key gap; and based on a preset angle constraint, determining the gap tilt threshold of the adjacent second-category keys according to the longitudinal contours of the adjacent second-category keys; when the tilt angle of the candidate inter-key gap meets the gap tilt threshold, using the corresponding candidate inter-key gap as the initial inter-key gap.

[0125] Specifically, based on the color features of the pixels in the first partial image, such as grayscale value or color saturation, narrow lines with sudden changes in brightness are detected as candidate key gaps. These candidate key gaps include not only real key gaps, but may also include pseudo gaps caused by reflections, stains or perspective deformation, such as oblique bright lines formed by reflections. The adjacent second-class piano keys adjacent to the candidate key gaps are obtained, such as the second-class piano keys on the left and / or right side of the candidate key gaps, and the second-class piano keys within a preset adjacent range of the candidate key gaps, wherein the perspective deformation degrees of the candidate key gaps and the adjacent second-class keys are related and similar. The actual arrangement direction of the longitudinal profiles of the adjacent second-class piano keys is analyzed, and the tolerable gap tilt threshold is calculated in combination with the preset angle constraint. Then, several candidate key gaps are screened, and the initial key gaps whose tilt degree meets the requirements are retained, while those exceeding the threshold are judged as noise or pseudo gaps and excluded.

[0126] The preset angle constraint is a preset tolerable angle deviation, and the specific value can be flexibly set according to actual needs, such as 15°, which is not limited here.

[0127] Among them, the longitudinal profile refers to the edge profile of the second-class piano keys extending along their long axis, such as the straight edges on both sides of the white keys, which is used to determine the relative angle between the main extension direction of the piano keys and the image coordinate system. It should be understood that since the longitudinal profile of the second-class piano keys may be tilted as a whole due to the pitch angle of the camera or the longitudinal profile may be tilted due to perspective, the gap tilt threshold is dynamically calculated according to the actual longitudinal profile direction, which can fit the key gap direction of the actual second-class piano keys in the image. For example, the actual longitudinal profile direction is tilted 5° to the right, combined with the preset angle constraint of 15°, then the gap tilt threshold is tilted 10° to the left to 20° to the right, and the candidate inter-key gaps with tilt angles within this range are used as the initial inter-key gaps.

[0128] Therefore, by combining the key contour direction and dynamic angle constraints, we ensure that the identified key gaps conform to the actual tilt angles of the keys in the image, and only retain the gaps between keys that conform to the perspective rules of specific image positions, effectively eliminating interference information and providing an accurate data basis for subsequent key gap repair and key positioning.

[0129] In some embodiments, a vertical edge enhancement strategy is employed based on the geometric characteristics of the parallel arrangement of the two types of keys. A customized horizontal gradient filter is used to prioritize the vertical gap features between the two types of keys while suppressing oblique interfering edges. The image gradient distribution is dynamically analyzed, and high and low thresholds are automatically set to ensure a balance between weak edge preservation and noise suppression. Simultaneously, the detected edges (i.e., candidate gaps between keys) are directional filtered, retaining only valid edges with a deviation of less than 15° from the direction of the two types of keys, and filtering for approximately vertical lines. This eliminates oblique interference such as shadows and reflections from the first type of keys through angular constraints, ensuring the physical consistency of the edges.

[0130] In some embodiments, the morphological operation includes the step of: S4031, obtaining at least one target second-class key adjacent to the initial inter-key gap to be processed, and updating the setting parameters of the structural element based on the perspective deformation parameters of the target second-class key.

[0131] Specifically, for each initial inter-key gap to be processed, or for multiple initial inter-key gaps to be processed within a preset common range, specific structural element setting parameters are used. By analyzing at least one target second-class piano key adjacent to the initial inter-key gap, the perspective deformation parameters of the target second-class piano key are obtained, and the setting parameters of the structural element used in the morphological operation are dynamically calculated and updated to match the actual deformation characteristics of the key gap. It should be understood that due to perspective, the keys are tilted or longitudinally contracted, resulting in deviations between the direction of the adjacent key gaps and the standard horizontal direction. At this time, the structural element is adaptively adjusted according to the specific perspective deformation of each key to perform differentiated processing on different key gaps.

[0132] In some embodiments, the difference in perspective deformation parameters of multiple initial inter-key gaps within a preset common range is smaller than the preset deformation difference, and the setting parameters of a structural element can be shared. The specific range division can be predetermined and is not limited here.

[0133] In some embodiments, the target second-class keys can be the second-class keys on the left and / or right side of the initial key gap. When multiple initial key gaps share the setting parameters of a structural element, the second-class keys on the left and / or right side of any one of the initial key gaps can be selected as the target second-class keys.

[0134] For example, the long axis direction of the structural element needs to be aligned with the direction of the adjacent target Class II keys so that the processed key gap matches the shape of the adjacent target Class II keys. The size parameters of the structural element can be adjusted according to the degree of key deformation, such as the length of the structural element shortens with perspective contraction.

[0135] It should be understood that the dynamic coupling mechanism based on perspective deformation parameters and structural element parameters enables morphological operations to accurately repair key cracks or eliminate noise, avoid processing deviations caused by fixed structural elements, and thus generate a target local image that is more consistent with the deformation laws of the piano keys in the image.

[0136] In some embodiments, the position coordinates of the target second-class piano key in the first field of view image are obtained, and the perspective deformation parameters of the target second-class piano key are determined based on preset camera parameters and the position coordinates of the target second-class piano key. For example, the position coordinates of the target second-class piano key in the first field of view image are obtained, and its perspective deformation parameters are calculated in combination with the preset camera parameters. For another example, for the same preset camera parameters, the perspective deformation parameters at each position coordinate are determined, and the corresponding perspective deformation parameters can be quickly obtained based on the preset camera parameters and the position coordinates of the target second-class piano key, wherein the perspective deformation parameters of each position coordinate in the image are determined based on prior data and are not limited here.

[0137] In some embodiments, the perspective deformation parameters include a shrinkage ratio and / or a shrinkage direction. The shrinkage ratio refers to the specific numerical ratio of the deformation of an object (e.g., a Class I piano key or a target Class II piano key) in the image due to perspective deformation, for example, the shrinkage ratio in the length direction or the shrinkage ratio in the width direction compared to an image without perspective deformation. The shrinkage direction refers to the direction in which the object (e.g., a Class I piano key or a target Class II piano key) is shortened in the image due to perspective deformation. For example, a distal piano key, as it is farther away from the camera, appears shortened in its longitudinal direction in the image, i.e., the shrinkage direction is the depth direction perpendicular to the line of sight.

[0138] In some embodiments, the setting parameters of the structural element include the major axis direction, the major axis length and the minor axis length, and the method also includes: obtaining the perspective deformation parameters of the target Class II keys, the perspective deformation parameters including the shrinkage ratio and / or the shrinkage direction; determining the major axis direction of the structural element based on the shrinkage direction; and / or determining the target morphological parameters of the target Class II keys according to the standard morphological parameters of the Class II keys and the shrinkage ratio; determining the major axis length and the minor axis length of the structural element according to the target morphological parameters of the target Class II keys.

[0139] Specifically, the long axis of the structuring element is dynamically adjusted based on the contraction direction of the target second-class keys. If a key shrinks longitudinally due to perspective, the contraction direction corresponds to the tilt of the key's long axis, and thus the tilt of the key seam. The long axis of the corresponding structuring element must be strictly aligned with this contraction direction to ensure that the morphological operation is consistent with the actual orientation of the key seam in the image. For example, if the far-end key appears to be shortening in the image and tilts upward and to the right, the long axis of the structuring element will rotate in the same direction.

[0140] Specifically, the standard morphological parameters of the second type of piano keys are obtained, such as the original length and width when no perspective deformation occurs, and the target morphological parameters are calculated in combination with the shrinkage ratio, such as the actual length and width after perspective shortening. Furthermore, the major axis length of the structural element is dynamically scaled and determined according to the actual length in the target morphological parameters, and / or, the ratio setting of the second type of piano keys to the key seam width is obtained, and the expected width of the key seam is determined according to the actual width and the ratio setting in the target morphological parameters. The minor axis length of the structural element is dynamically scaled and determined according to the expected width of the key seam, ensuring that the morphology of the structural element is highly matched with the physical characteristics of the actual key seam. Thus, by combining the perspective deformation parameters with the standard morphological parameters, the size parameters of the structural element are accurately adapted to the key shape under the current viewing angle, thereby improving the accuracy of key seam repair and noise suppression.

[0141] In some embodiments, the morphological operation further includes step: S4032, performing a morphological operation on the initial inter-bond gaps to be processed based on the updated structural elements.

[0142] Specifically, the updated structuring element is used to perform morphological operations on one or multiple initial inter-key gaps within a preset common range to optimize the gap morphology. Through dynamic adaptive morphological processing, the structuring element can accurately fit the actual key gap direction and deformation pattern, retaining effective features while suppressing noise. As shown in 40d in Figure 11, the morphological operation thickens, expands, and inflates the white key edges (i.e., the initial inter-key gaps), ultimately generating a target local image containing only complete, continuous, and physically compliant inter-key gaps, thereby improving the robustness and accuracy of subsequent key positioning.

[0143] In some embodiments, the method further includes: based on the updated structural element, expanding the boundary of the initial gap between keys to fill the internal holes or breaks of the initial gap between keys to obtain a second local image; based on the updated structural element, shrinking the boundary of the initial gap between keys in the second local image to eliminate the burr noise and linear noise of the initial gap between keys to obtain the target local image.

[0144] Specifically, the key gap features are optimized through a two-stage morphological operation. First, based on the updated structuring element, an expansion operation, such as morphological dilation, is performed on the boundary of the initial key gap. The sliding of the structuring element fills the internal holes or broken areas in the key gap, generating a second local image and ensuring the continuity of the key gap. Second, the structuring element is again used to perform a contraction operation, such as morphological erosion, on the key gap boundary of the second local image. This eliminates edge expansion or pseudo-key gaps caused by over-expansion, ultimately forming a target local image that only retains the complete, smooth, and physically correct key gap.

[0145] In some embodiments, the initial inter-key gaps are traversed based on the updated structural element to expand the set of pixels that match the structural element and / or shrink the set of pixels that do not match the structural element. For the pixel set that matches the morphology of the structural element, such as the pixel set whose direction is adapted to the key gap tilt angle and whose size matches the perspective scaling ratio, an expansion operation is performed to repair key gap breakage and key sticking caused by image noise or viewing angle tilt; and for the pixel set that does not match the structural element, such as the pixel set whose direction is not adapted to the key gap tilt angle, a contraction operation is performed to eliminate redundant areas or areas that deviate from the expected morphology, such as pseudo gaps formed by edge glitches or background interference.

[0146] In some embodiments, the structuring element is elliptical and, as the direction of its major axis changes, exhibits a tilted elliptical form, i.e., a tilted elliptical kernel. Exemplarily, the tilted elliptical kernel is an asymmetric, directionally adjustable morphological structuring element. It is elliptical in shape, and the lengths of its major and minor axes are dynamically calculated based on the actual physical dimensions of the second-class piano keys and camera parameters. The major axis direction aligns with the perspective contraction direction of the second-class piano keys, enhancing the continuity of the key gap. The kernel parameters vary with image position to match the degree of deformation in different regions.

[0147] In some embodiments, a mathematical model of the kernel size is constructed based on the actual physical size of the second type of piano keys (approximately 23 mm wide) and preset camera parameters (such as focal length and tilt angle). In the edge deformation area of ​​the image, a tilted elliptical kernel is used for gap connection, and its long axis direction is consistent with the perspective contraction direction of the white keys to achieve geometric compensation.

[0148] It should be understood that the closer the keys are to the sides of the image, the smaller the key gaps are, and under high light intensity, two white keys may become stuck together. The essential process of morphological analysis is expansion and erosion. In the embodiments of this application, the kernel that determines the degree of expansion and erosion is dynamically updated to avoid using the same kernel for processing, thereby preventing the kernel from being too large, thereby reducing the white key judgment area, and the kernel from being too small, thereby failing to separate stuck white keys.

[0149] It should be understood that the introduction of a coupling mechanism between dynamic structuring elements and perspective deformation parameters allows the morphological parameters of the structuring elements (such as the major axis direction and the length of the major and minor axes) to be adjusted in real time according to perspective deformation parameters such as the tilt angle and contraction direction of the keys, so that they accurately match the direction and morphological changes of the key seams, effectively improving the accuracy and anti-interference performance of key recognition. For example, if the target second-class piano key has a key seam direction that deviates from the vertical direction due to camera tilt, the major axis direction of the structuring element is adjusted to be consistent with the key seam tilt angle, ensuring that the morphological operation is performed along the direction of the key seam deformation and accurately repairing breakage or peeling noise; if the key seam width is scaled due to perspective projection, the size of the structuring element is reduced according to the scaling ratio to avoid overfilling or erosion, allowing the structuring element to adapt to the key seam deformation under different shooting conditions, taking into account the consistency of the geometric features of the near and far key seams, thereby improving the robustness of the morphological operation.

[0150] S404: Identify a plurality of the second-category keys from the target partial image based on a third preset model.

[0151] Specifically, by inputting the complete, continuous and noise-free gaps between keys in the target local image, the third preset model is provided with clear key boundary features. The third preset model combines the pre-trained feature library (such as key shape, size, arrangement pattern) or rule algorithm (such as contour matching, proportion analysis) to accurately locate and classify the second type of keys that meet the preset features. For example, if the target is a white key, a wide continuous bright area is extracted. Figure 11 As shown in 40e, the key image areas of several second-class keys are extracted from the target local image and compared with the first local image, where the red lines are the extracted second-class key contours. It can be seen that the third preset model accurately extracts each second-class key without omission.

[0152] It should be understood that because the key seams undergo dynamic structural elements to repair breaks, eliminate burrs, and strictly match the perspective deformation of the keys, the key seam boundary contours exhibit a high degree of standardization, which means that the boundary contours of the keys also exhibit a high degree of standardization. The third preset model can identify the second class of keys from the target partial image simply through simple rules or pre-trained feature matching. Therefore, the third preset model is a pre-configured recognition algorithm or a trained classification model, whose feature library or rules are constructed based on prior information such as the geometry and arrangement of the first class of keys. Existing algorithms or contour matching models can also be used, and this is not limited here.

[0153] In some embodiments, the method further includes: obtaining at least one proximal inter-key gap whose perspective deformation parameter is lower than a first deformation threshold, and determining a baseline morphological parameter of the inter-key gap based on the morphological parameters of at least one proximal inter-key gap; obtaining at least one distal inter-key gap whose perspective deformation parameter is higher than a second deformation threshold, wherein the first angle threshold is less than or equal to the second angle threshold; adjusting the baseline morphological parameter of the inter-key gap based on the perspective deformation parameter of each distal inter-key gap, and determining the target morphological parameter of each distal inter-key gap; and adjusting the contour of the corresponding distal inter-key gap according to the target morphological parameter of each distal inter-key gap.

[0154] Specifically, near-end inter-key gaps (such as the gaps between the second-class piano keys directly in front of the camera with almost no perspective distortion) with perspective deformation parameters lower than the first deformation threshold are screened out, and their morphology is close to the reference morphology under the standard viewing angle. By counting the morphological parameters of at least one near-end inter-key gap, the reference morphological parameters of the inter-key gaps are calculated and updated as a reference for subsequent adjustments. For far-end inter-key gaps (such as the gaps between the second-class piano keys with obvious deformation due to tilt at the edge of the picture) with perspective deformation parameters higher than the second deformation threshold, the reference morphological parameters are dynamically adjusted through perspective transformation based on their specific perspective deformation parameters. For example, if the length of the far-end key gap is shortened due to perspective, the reference length is scaled proportionally; if the direction is offset, the long axis direction is adjusted to match the tilt angle.

[0155] Based on the adjusted target morphological parameters, the contour of the distal key seam is adjusted, such as Figure 11 As shown in Figure 40f, the green line represents the adjusted second-class key outline. This outline accurately fits the second-class key while maintaining a smooth, standardized shape. This ensures that its shape is consistent with the baseline features of the proximal key seam, improving the geometric standardization and recognition accuracy of the overall key outline. It also adapts to the deformation patterns caused by the viewing angle, allowing the subsequently extracted second-class keys to also adapt to the deformation patterns caused by the viewing angle.

[0156] In this embodiment, a key type with distinct features is prioritized for identification. Based on the spatial correlation between the keys (e.g., shape or area ratio), the local image of the second key type to be identified is quickly locked onto, converging the global field of view to the local image. This significantly reduces the image processing scope, computational complexity, and background interference while ensuring positioning accuracy. The natural boundary between the keys—the key seam—is then selected as the core processing unit. A reverse normalization strategy is employed to avoid directly processing large, homogeneous areas. This further reduces the image processing scope to the microstructure of the key seam, further reducing computational complexity.

[0157] In some embodiments, the capture of the first partial image is preferably performed before the first filtering operation and edge detection (i.e., initial key gap detection) to avoid unnecessary signal-to-noise ratio degradation. Secondly, morphological operations are preferably performed after edge detection, using edge intensity maps to enhance feature connectivity and reduce the impact of isolated noise, thereby avoiding an increase in the key breakage rate. Therefore, a progressive path of edge detection, morphological operations, and contour extraction (i.e., identification of the second type of keys) is adopted to gradually enhance features. Finally, physical verification (such as spacing checks and aspect ratio filtering) is performed at the end of the process. At this time, judgments are made based on complete contour information to prevent premature pruning and accidental deletion of valid contours due to local deformation or information fragmentation.

[0158] In some embodiments, see Figure 12 , Figure 12 is a schematic diagram of a piano key recognition result provided by an embodiment of the present application, such as Figure 12 As shown, a piano key recognition method provided in an embodiment of the present application can be used to identify Class I keys (such as the black keys in the figure), with the recognition results shown in red. A key seam-based piano key recognition method provided in an embodiment of the present application can be used to identify Class II keys (such as the white keys in the figure), with the recognition results shown in green. Furthermore, the recognition results can be the position coordinates, key outlines, and key image regions of several Class I keys and several Class II keys.

[0159] In some embodiments, in the step of determining the target morphological parameters of each of the keys of the first category based on the baseline morphological parameters of the keys and the perspective deformation parameters, and adjusting the outline of the corresponding keys of the first category based on the target morphological parameters of each of the keys of the first category, the top surface position of the keys of the first category is determined based on the position coordinates of the keys of the first category in the image, and based on the top surface position, the outline of the corresponding keys of the first category is adjusted according to the target morphological parameters of each of the keys of the first category, so that the outline of the corresponding keys of the first category includes the top surface position. It should be understood that the three-dimensional height difference of the black keys causes perspective deformation and edge occlusion effects during imaging, which in turn causes the top surface to shift, such as Figure 6 As shown in the figure, the black key on the far right has the problem of top surface offset. At this time, according to its position coordinates, its top surface position is predicted to be on the right side of the outline, and the outline is adjusted with the right side of the outline as the limit boundary. That is, if the outline needs to be reduced (such as adjusting the outline according to the expected width), the left side of the outline will be reduced first, while the right part of the image will be retained, as shown in the figure. Figure 12 As shown in the figure, the adjusted outline prioritizes the right image, ensuring accurate selection of the key top surface. This provides reliable data for subsequent positioning or interactive operations, improving adaptability in complex technique analysis, teaching feedback, and smart instrument interaction.

[0160] Exemplarily, the above method can be implemented in the form of a computer program that can be run on an intelligent musical instrument. The intelligent musical instrument includes keys for directly interacting with a performer and playing music; an image acquisition module for acquiring a first field of view image, wherein the first field of view image includes images corresponding to a number of first-class keys or the first field of view image includes images corresponding to a number of first-class keys and second-class keys. The intelligent musical instrument also includes a memory for storing a computer program; a processor for executing the computer program and implementing the key recognition method and / or the key seam-based key recognition method provided in any embodiment of the present application when executing the computer program.

[0161] For example, see Figure 13 , Figure 13 is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The above method can be implemented in the form of a computer program, which can be used in Figure 13 Run on the computer device shown. Figure 13 As shown, the computer device includes a processor, a memory and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium may store an operating system and a computer program. The computer program includes program instructions, which, when executed, may enable the processor to execute any one of the key recognition methods and / or the key recognition method based on key seams. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium, which, when executed by the processor, may enable the processor to execute any one of the key recognition methods and / or the key recognition method based on key seams. The network interface is used for network communication, such as sending assigned tasks, etc.

[0162] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0163] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0164] S301, acquiring a first field of view image, wherein the first field of view image includes images corresponding to several first-class piano keys; S302, performing a first filtering operation on the pixel points in the first field of view image based on the pixel values ​​in the first field of view image to suppress the first form of image noise, and obtaining a second field of view image; S303, performing preset threshold processing on the second field of view image to obtain a third field of view image, wherein the third field of view image includes several target pixel sets, and the target pixel sets are composed of several continuous pixel points whose pixel values ​​are lower than the preset threshold; S304, performing a second filtering operation on the target pixel sets in the third field of view image based on a preset structural element to suppress the second form of image noise, and obtaining a fifth field of view image; S305, identifying several first-class piano keys from the fifth field of view image based on a second preset model.

[0165] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0166] S401, obtaining a first field of view image, the first field of view image including images corresponding to several first-class keys and second-class keys; S402, based on the preset positional relationship between the first-class keys and the second-class keys, and according to the key image area where the first-class keys are located, intercepting a first partial image of the area where several second-class keys are located from the first field of view image; S403, identifying several initial inter-key gaps between adjacent second-class keys in the first partial image; and performing morphological operations on the several initial inter-key gaps in the first partial image to obtain a target partial image; S404, identifying several second-class keys from the target partial image based on a third preset model; wherein the morphological operation includes the steps of: S4031, obtaining at least one target second-class key adjacent to the initial inter-key gap to be processed, and updating the setting parameters of the structural element based on the perspective deformation parameters of the target second-class key; S4032, performing morphological operations on the initial inter-key gap to be processed based on the updated structural element.

[0167] Exemplarily, the processor is used to run a computer program stored in the memory, and is also used to implement the steps of the key recognition method and / or the key seam-based key recognition method provided in any embodiment of the present application, which will not be repeated here.

[0168] A computer-readable storage medium is also provided in an embodiment of the present application, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the processor executes the program instructions to implement the steps of any one of the key recognition methods and / or key seam-based key recognition methods provided in the embodiments of the present application.

[0169] The computer-readable storage medium may be an internal storage unit of the intelligent musical instrument described in the aforementioned embodiment, such as a hard disk or memory of the intelligent musical instrument. The computer-readable storage medium may also be an external storage device of the intelligent musical instrument, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the intelligent musical instrument.

[0170] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for identifying a piano key, characterized in that: The method comprises: S301, acquiring a first field of view image, wherein the first field of view image includes images corresponding to a plurality of piano keys of a type; S302: Performing a first filtering operation on pixels in the first field of view image based on pixel values ​​in the first field of view image to suppress a first form of image noise, thereby obtaining a second field of view image; wherein the first filtering operation is based on the grayscale or color distribution characteristics of each pixel in the first field of view image to adaptively smooth the local area; and the first form of image noise refers to discrete, isolated interference present in the image; S303, performing preset threshold processing on the second field of view image to obtain a third field of view image, wherein the third field of view image includes a plurality of target pixel sets, each of which is composed of a plurality of consecutive pixel points whose pixel values ​​are lower than a preset threshold; S304, performing a second filtering operation on the target pixel set in the third field of view image based on a preset structuring element to suppress image noise of the second form, thereby obtaining a fifth field of view image; wherein the preset structuring element is a preset set of pixels used to detect and correct shape defects of the target pixel set; and the second filtering operation is a morphological operation performed based on the preset structuring element; S305 : Based on the second preset model, identify a plurality of keys of the type from the fifth field of view image.

2. The method according to claim 1, characterized in that The S302 includes: Identifying at least one pixel curve on the first field of view image, where the pixel curve is composed of a plurality of pixel points and is used to reflect a degree of change in pixel values ​​of the pixel points; Calculating a pixel difference between at least one first target segment and an adjacent second target segment in the pixel curve, wherein the first target segment and the second target segment are composed of at least one pixel point; Determining whether the degree of pixel difference exceeds a pixel difference threshold; If so, the first target segment is corrected to reduce the pixel difference.

3. The method according to claim 1, characterized in that The second form of image noise includes burr noise and linear noise; S304 includes: Based on the first structuring element, expanding the boundaries of each target pixel set in the third field of view image to fill internal holes or breaks in the target pixel set to obtain a fourth field of view image; Based on the second structure element, the boundaries of each target pixel set in the fourth field of view image are shrunk to eliminate burr noise of each target pixel set and linear noise in the fourth field of view image, thereby obtaining the fifth field of view image.

4. The method according to claim 1, wherein The method further comprises: Based on preset camera parameters, determining the perspective deformation parameter of each of the keys of the type according to the position coordinates of the keys of the type in the fifth field of view image; The target shape parameters of each key of the type are determined according to the reference shape parameters of the key of the type and the perspective deformation parameters, and the contour of the corresponding key of the type is adjusted according to the target shape parameters of each key of the type.

5. The method according to claim 4, characterized in that The method further comprises: Acquire at least one proximal key of the first type whose perspective deformation parameter is lower than a first deformation threshold, and determine reference morphological parameters of the first type of keys based on the morphological parameters of the at least one proximal key; Acquire at least one distal key of the first type having the perspective deformation parameter higher than a second deformation threshold, wherein the first deformation threshold is less than or equal to the second deformation threshold; Adjusting the reference standard morphological parameters of the first type of keys based on the perspective deformation parameters of each of the first type of remote keys to determine the target morphological parameters of each of the first type of remote keys; According to the target morphological parameters of each distal key, the contour of the corresponding distal key is adjusted.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Obtain the actual shape parameters and position coordinate parameters of each type of piano key; Based on the preset key layout parameters, a plurality of keys of the first category are screened according to the actual shape parameters and position coordinate parameters of the keys of the first category to determine a plurality of target keys of the first category; Output the position coordinates of a plurality of the target first-class piano keys in the first field of view image.

7. The method according to claim 6, characterized in that The preset key layout parameters include at least one of a target length, a target width, a target aspect ratio, a target area, a target key spacing, and a target number of keys for the type of keys.

8. An intelligent musical instrument, characterized in that: The intelligent musical instrument comprises: An image acquisition module, configured to acquire a first field of view image, wherein the first field of view image includes images corresponding to a plurality of piano keys of a type; memory for storing computer programs; A processor is configured to execute the computer program and implement the key recognition method according to any one of claims 1 to 7 when executing the computer program.

9. A computer device, characterized in that: The device comprises: memory for storing computer programs; A processor is configured to execute the computer program and implement the key recognition method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the key recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent playing error identification method and system for assisting piano teaching

    CN113723264A

  • Key identification method and device, electronic equipment and storage medium

    CN111695499A

  • Intelligent identification method and system for giving assistance with piano teaching, and intelligent piano training method and system

    WO2022052941A1