Key identification method, intelligent musical instrument, equipment and medium
Through phased and gradual interference filtering mechanism and perspective deformation correction, the problem of insufficient accuracy and robustness of key recognition in complex scenarios is solved, and high-precision key recognition is achieved.
Patent Information
- Application Number
- CN202510585626.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing key recognition technology lacks recognition accuracy and robustness in complex scenarios, and is greatly affected by natural light, stage light, moving objects and environmental background.
A staged and gradual interference filtering mechanism under strict timing is adopted to gradually purify the image through multi-directional smoothing strategy, preset threshold processing and morphological operations, and combine perspective deformation correction and physical verification to achieve high-precision recognition of the keys.
In complex scenarios, the recognition accuracy and robustness of the keys are significantly improved, ensuring the retention of key details and effective removal of interference.
Smart Images

Figure CN120125810A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital performance, and particularly to a key recognition method, intelligent musical instrument, device, and medium. Background Art
[0002] With the development of technology, the digital expression forms of music and art are becoming increasingly rich, and have broad application prospects in fields such as education, entertainment, and virtual reality. Among them, key recognition technology is the core technology in fields such as music teaching, performance assistance, and intelligent musical instrument development. Its core goal is to detect the position and state of keys through sensors or vision systems.
[0003] For example, the invention patent application with the publication number CN111695499A discloses a key recognition method, device, electronic device, and storage medium. The method includes: collecting a first keyboard image of a target musical instrument, where black keys and white keys are arranged on the keyboard of the target musical instrument; inputting the first keyboard image into a preset image recognition model to determine the contours of the black keys in the first keyboard image; and determining the key range corresponding to each octave and the contours of the white keys in each octave in the first keyboard image according to the contours of the black keys and the spacing between adjacent black keys.
[0004] Another example is the Chinese invention patent application with the publication number CN113723264A, which discloses a method for intelligently identifying playing errors for assisting piano teaching, including: obtaining a 2D image of a piano being played that includes a complete piano keyboard from above the piano keyboard; performing object detection on the 2D image through a piano keyboard detection network to detect the piano keyboard area represented by the relative position coordinates of the 2D image, and obtaining the piano keyboard position coordinates in the original coordinates of the 2D image through conversion.
[0005] However, in practical applications, environmental noises such as natural light or stage lighting and moving objects will affect the quality of the key images captured by the vision system, thereby leading to a decrease in key recognition accuracy and insufficient robustness in complex scenarios. Summary of the Invention
[0006] The main purpose of this application is to provide a key recognition method, intelligent musical instrument, device, and medium. To solve the above-mentioned technical problems, this application specifically adopts the following technical solutions: In the first aspect of this application, a key recognition method is provided. The method includes: S301, obtaining a first field of view image, where the first field of view image includes images corresponding to a plurality of first-type keys; S302, perform a first filtering operation on the pixel points in the first field of view image based on the pixel values in the first field of view image to suppress image noise of a first form, and obtain a second field of view image; S303, perform a preset threshold processing on the second field of view image to obtain a third field of view image, where the third field of view image includes a number of target pixel sets, and each target pixel set is composed of a number of consecutive pixel points with pixel values lower than the preset threshold; S304, perform a second filtering operation on the target pixel sets in the third field of view image based on a preset structural element to suppress image noise of a second form, and obtain a fifth field of view image; S305, identify a number of the first type of keys from the fifth field of view image based on a second preset model.
[0007] In a second aspect of the present application, there is provided an intelligent musical instrument, and the intelligent musical instrument includes: An image acquisition module, configured to acquire a first field of view image, where the first field of view image includes images corresponding to a number of the first type of keys; A memory, configured to store a computer program; A processor, configured to execute the computer program and implement the steps of the key recognition method provided in any embodiment of the present application when executing the computer program.
[0008] In a third aspect of the present application, there is provided a computer device, and the device includes: A memory, configured to store a computer program; A processor, configured to execute the computer program and implement the steps of the key recognition method provided in any embodiment of the present application when executing the computer program.
[0009] In a fourth aspect of the present application, there is correspondingly provided a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the steps of the key recognition method provided in any embodiment of the present application.
[0010] Beneficial effects: The present application provides a key recognition method, an intelligent musical instrument, a device, and a medium, and specifically provides a phased and progressive interference filtering mechanism under strict timing. The filtering intensity in each stage is strictly limited, and only specific interference information is processed. In the early stage, the original image data is highly respected for fine-tuning, and in the later stage, adaptive correction is performed based on the reliable image data and camera parameters after processing. During the process, some interference information is tolerated to be passed backward and then gradually eliminated, avoiding the loss of details caused by one-time highly sensitive filtering. Thus, while retaining the details of the keys, high-precision interference elimination is achieved, effectively improving the recognition accuracy and robustness of the keys in complex scenarios.
[0011] In the first stage, only local smoothing is performed on discrete isolated noise points (such as sensor noise or dust reflection). The multi-directional smoothing strategy based on the pixel distribution characteristics can keep the overall structure of the image unchanged and avoid the loss of details caused by excessive denoising. The edge blurring that may be caused by the smoothing strategy is clarified and restored during threshold processing to ensure the geometric integrity of the keys, thereby providing a reliable parameter calibration basis for the second stage.
[0012] In the second stage, through conservative parameter setting for pixel aggregation, the microscopic fractures generated inside the black keys due to the shooting angle (such as the microscopic fractures inside the black keys caused by wood texture or uneven illumination) are filled, and a complete closed area is reconstructed. Then, refined boundary trimming is implemented to synchronously eliminate structural noise (such as the burrs on the edges of the keys, the linear interference formed by the key seams and reflections on the keyboard) and background interference, realizing the synchronous optimization of repairing effective features and stripping invalid features, and avoiding the generation of multi-segment broken edges in the unclosed contours during the subsequent extraction process, which may lead to confusion in the contour hierarchical relationship.
[0013] In the third stage, based on perspective deformation, the key contours are adaptively corrected. Through the sub-region processing and dynamic parameter correction of the keys proximal to and distal from the optical axis, it is ensured that the calibrated key contours restore the deformation errors caused by the shooting angle, so as to accurately locate the keys at various positions in the image. Even when the instrument model and shooting angle change, a high recognition rate can still be maintained, and it has stronger generalization ability.
[0014] In the fourth stage, a hard physical verification of the key layout parameters is performed. In order to avoid premature pruning that may cause valid contours to be mistakenly deleted due to local deformation, the physical verification is carried out at the end of the process. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts do not necessarily draw according to the actual ratio. Obviously, the following-described drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0016] Figure 1 is a schematic flowchart of a key recognition method provided by an embodiment of the present application; Figure 2 is a schematic diagram of a first field of view image provided by an embodiment of the present application; Figure 3 is a schematic diagram of a second field of view image provided by an embodiment of the present application; Figure 4 It is a schematic diagram of a third field of view image provided by an embodiment of the present application; Figure 5 It is a schematic diagram of a fifth field of view image provided by an embodiment of the present application; Figure 6 It is a schematic diagram of the extraction of a first type of key contour provided by an embodiment of the present application; Figure 7 It is a schematic diagram of another extraction of a first type of key provided by an embodiment of the present application; Figure 8 It is a schematic diagram of a first type of key contour after adjustment provided by an embodiment of the present application; Figure 9 It is a schematic flow chart of a key recognition method based on key gaps provided by an embodiment of the present application; Figure 10 It is a schematic diagram of intercepting a first partial image provided by an embodiment of the present application; Figure 11 It is a schematic diagram of recognizing a key based on key gaps provided by an embodiment of the present application; Figure 12 It is a schematic diagram of a key recognition result provided by an embodiment of the present application; Figure 13 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0018] In this article, suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of description of the present application, and they have no specific meaning in themselves. Therefore, "module", "component", or "unit" can be used interchangeably.
[0019] In this text, terms such as "upper", "lower", "inner", "outer", "front", "rear", "one end", "the other end", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of this application. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0020] In this text, unless otherwise clearly defined and limited, terms such as "installed", "provided with", "connected", etc. shall be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0021] In this text, "and / or" includes any and all combinations of one or more of the listed related items.
[0022] In this text, "a plurality of" means two or more, that is, it includes two, three, four, five, etc.
[0023] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising that element.
[0024] In practical applications, the environment where the keys are located has a high degree of variability, and the image acquisition module can be diversely set at different positions. The current key recognition technology has low environmental adaptability and is difficult to maintain high reliability in practical applications, which limits its wide application. On the one hand, the changes in natural light or stage lighting introduce additional noise, which can cause the images captured by the image acquisition module to be overexposed, reflective or color distorted, reducing the accuracy of feature extraction. And the complex environmental background (such as moving objects, cluttered background) causes the target to be mixed with the background, increasing the recognition difficulty and significantly decreasing the recognition accuracy. On the other hand, different setting positions of the image acquisition module will cause changes in the viewing angle, which in turn makes the keys present different states in the image, further increasing the recognition difficulty.
[0025] Based on this, the present application proposes a key recognition method, specifically providing a phased and progressive interference filtering mechanism under strict timing. The filtering intensity of each stage is strictly limited, and it only processes specific interference information. In the early stage, it highly respects the original image data for fine-tuning, and in the later stage, it performs adaptive correction based on the reliable image data and camera parameters after processing. During the process, it tolerates some noise or blurred information to be transmitted to the subsequent stage, gradually eliminating various interference information in the image, avoiding the loss of details caused by one-time highly sensitive filtering. Thus, while retaining the details of the keys, high-precision interference removal is achieved, effectively improving the recognition accuracy and robustness of the keys in complex scenarios.
[0026] Furthermore, the present application also proposes a key recognition method based on key slots. Taking a type of keys with prominent features as a guide, it quickly locks the local range of the second type of keys to be recognized, normalizes the key slots with prominent features in the local range based on the deformation compensation of the dynamic structural element, and then deduces the regional contour of the keys through the key slots in reverse. Thus, through the multi-level focus area convergence strategy of "global image - local image - key - to - key gap", combined with the coupling mechanism of the dynamic structural element and the perspective deformation parameter, the dual optimization of the efficient utilization of computing resources and the recognition accuracy is achieved, and the anti-interference ability of the recognition is effectively improved.
[0027] In this article, the musical instrument has keys with a fixed arrangement order, and each key has a fixed pitch device. These keys can form a keyboard. The musical instrument can specifically be a piano, an organ, an accordion, an electronic keyboard, etc., which is not limited here.
[0028] In this article, the image acquisition module is a device that converts the optical signal in the physical world into a digital image, such as a camera, a video camera, etc. Its core components can include an optoelectronic converter (such as a CCD or CMOS sensor), a lens, a signal processing unit, and a storage module. The model and specific components of the image acquisition module can be selected according to actual needs, which is not limited here.
[0029] Exemplarily, the installation position of the image acquisition module can be determined according to the specific application scenario (such as ambient light), technical parameters (such as resolution, installation height, angle, picture integrity, etc.) to ensure the device performance and compliance. The specific installation position is not limited here. For example, the image acquisition module can be perpendicular to the key plane and located directly above the keyboard, such as inside the piano lid, on top of the piano stand, or on a hanging bracket; or for another example, the image acquisition module can form a certain angle with the key plane and be installed on the front side or the upper side of the piano. Or for another example, the image acquisition module is independently set at the position corresponding to the user's perspective.
[0030] In this text, perspective is a fundamental law in optical imaging, which describes the deformation of three-dimensional objects presenting "larger near and smaller far" in a two-dimensional image due to distance differences. Correspondingly, the perspective deformation parameters reflect the deformation characteristics of the keys caused by the camera's perspective, such as the contraction ratio, contraction direction, etc. The position closer to the image acquisition module has a lower degree of perspective than the edge position, and the corresponding perspective deformation parameters also change accordingly.
[0031] It should be understood that the image acquisition module has preset internal parameters and external parameters, that is, preset camera parameters, including focal length, installation height, pitch angle, horizontal deflection angle, sensor size, lens distortion coefficient, etc., which can be used to calibrate the position of the optical axis. According to these preset camera parameters and the position of the object in the image (such as the position coordinates of the first-class keys or the second-class keys), the perspective deformation parameters of the object are determined. For example, according to the included angle parameter (i.e., the perspective angle) between the position of the object in the image and the optical axis of the image acquisition module, the perspective deformation parameters of the object are determined. The position closer to the optical axis has a smaller perspective angle and a lower corresponding deformation degree than the edge position.
[0032] In this text, morphological operation is a set of mathematical methods for processing digital images, mainly used to extract specific shape information of image components. It is based on the principle of set theory, and the main operations include dilation, erosion, opening operation, and closing operation, etc., to achieve functions such as noise removal, shape feature enhancement or weakening. And the structuring element is the core tool for image processing in mathematical morphology. Its essence is a set of pixel points with known shape, size, and direction, which defines the shape and scope of action of the morphological operation.
[0033] The following will describe some embodiments of the present application in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other. Please refer to Figure 1 , Figure 1 is a schematic flowchart of a key recognition method provided by an embodiment of the present application. As Figure 1 shown, an embodiment of the present application provides a key recognition method; the method includes S301 to S305.
[0034] S301, obtain a first field of view image, where the first field of view image includes images corresponding to a number of first-class keys.
[0035] Among them, the first field of view image refers to the original image directly collected by the image acquisition module, and this image needs to include the first-class keys to be recognized as the input basis for subsequent key recognition.
[0036] Please refer to Figure 2 , Figure 2 is a schematic diagram of a first field of view image provided by an embodiment of the present application. As Figure 2As shown, the musical instrument is a piano, and the image acquisition module is vertically arranged on the key plane and directly above the keyboard. The first field of view image includes several first-class keys 100 to be recognized (such as the black keys in the figure), and may also include other components of the musical instrument, such as the shell, strings, hammers, etc., as well as the venue background where the musical instrument is located and the hand movements of the performer. It can be seen that in the actual performance scenario, there are complex interfering environmental factors, such as Figure 2 key reflection in Figure 2
[0037] S302. Perform a first filtering operation on the pixel points in the first field of view image based on the pixel values in the first field of view image to suppress the image noise of the first form, and obtain a second field of view image.
[0038] Specifically, the image noise of the first form refers to the discrete and isolated interferences existing in the image, such as random discrete noise points generated by the sensor, dust reflection, or tiny light spots in the environment. The performance of this type of noise in the image is interference points with significant differences from the surrounding pixel values but scattered spatial distributions, which may interfere with the integrity and continuity of key structures such as the key contour.
[0039] Correspondingly, the first filtering operation is an image processing method for adaptively smoothing a local area based on the gray or color distribution characteristics of each pixel point in the first field of view image. For example, by analyzing the pixel value distribution within the pixel neighborhood, a multi-directional smoothing strategy is used to process the first field of view image, thereby suppressing the image noise of the first form.
[0040] In some embodiments, only selective smoothing is performed on the first field of view image. For example, for pixels suspected of being noise points (such as those with a difference from the neighborhood median exceeding a preset median threshold), they are replaced with neighborhood statistical values (such as the median or weighted mean), while the original values are retained for the key edges.
[0041] In some embodiments, the multi-directional smoothing strategy can be median filtering, Gaussian filtering, or adaptive weight filtering. Preferably, the multi-directional smoothing strategy is Gaussian filtering. It should be understood that Gaussian filtering assigns weights using a Gaussian function, with pixels closer to the center being given higher weights and pixels farther from the center having lower weights. This weight assignment mechanism makes Gaussian filtering more natural and smooth when removing noise, but it also cannot completely eliminate the noise. Therefore, sacrificing some denoising effects to ensure the retention of image details and key edges.
[0042] It should be understood that in the first stage, only local smoothing is performed on discrete isolated noise points (such as sensor noise or dust reflection). The multi-directional smoothing strategy based on the pixel distribution characteristics can keep the overall structure of the image unchanged, avoid losing details due to excessive denoising, ensure the geometric integrity of the keys, and thus provide a reliable parameter calibration basis for the second stage. Since the interference filtering intensity in the first stage is strictly limited, after the first stage of processing, the generated second field of view image only preliminarily purifies the discrete noise. While retaining the clarity of the main structure of the keys, it also retains complex environmental interferences such as local breaks caused by shooting angles, uneven lighting, or black key textures. Please refer to Figure 3 , Figure 3 which is a schematic diagram of a second field of view image provided by an embodiment of the present application. As Figure 3 shown, macroscopically, there are still environmental interferences such as key reflections and cluttered backgrounds in the second field of view image.
[0043] In some embodiments, S302 includes: identifying at least one pixel curve on the first field of view image, the pixel curve being composed of a plurality of pixel points, and the pixel curve being used to reflect the change degree of the pixel values of the pixel points; calculating the pixel difference degree between at least one first target segment and an adjacent second target segment in the pixel curve, wherein the first target segment and the second target segment are composed of at least one pixel point; determining whether the pixel difference degree exceeds a pixel difference threshold; if so, correcting the first target segment to reduce the pixel difference degree.
[0044] Specifically, at least one pixel curve in the first field of view image is identified. The pixel curve is composed of continuously arranged pixel points, and its pixel values show dynamic changes along the curve path, such as light and dark transitions or texture trends. Further, the pixel curve can be segmented with a fixed length or adaptively to obtain several target segments of continuous pixel points. For any first target segment, at least one adjacent second target segment is obtained, and the pixel difference degree of the target segments is calculated, such as mean difference, variance ratio, or maximum gradient difference, etc. If the pixel difference degree exceeds the pixel difference threshold, a noise area presenting as a mutated abnormal peak or valley is identified, and the first target segment is corrected, for example, replaced with a smooth transition value or interpolated reconstruction of the second target segment to eliminate the mutation and reduce the difference degree. Among them, the pixel difference threshold is a preset difference tolerance critical value used to trigger the correction operation, and the specific value can be determined according to noise feature statistics or empirical values, which is not limited here.
[0045] In some embodiments, if not, the original pixel values of the first target segment are retained to avoid redundant correction of normal pixels and maintain the integrity of the key edges and details.
[0046] In some embodiments, multiple pixel curves are generated in a specific direction in the first field of view image, which may specifically be the horizontal direction, the vertical direction, the tangent direction of the edge of the key, etc., and are not limited herein.
[0047] Thus, by analyzing the pixel value differences between adjacent regions segment by segment, the target segments showing continuous or step changes are distinguished, the first form of noise is accurately located and suppressed, and the overcorrection that may cause the key contour to be blurred is avoided, providing a clearer intermediate image for subsequent threshold processing.
[0048] S303. Perform a preset threshold processing on the second field of view image to obtain a third field of view image, where the third field of view image includes a plurality of target pixel sets, and each target pixel set is composed of a plurality of consecutive pixel points with pixel values lower than the preset threshold.
[0049] Among them, the preset threshold processing is a processing step of converting the grayscale or color information of the second field of view image into a binary image, and the image pixels are divided into foreground and background by setting a fixed threshold or an adaptively calculated adaptive threshold. Preferably, the preset threshold processing is an adaptive binarization corresponding to the adaptive threshold, and the threshold is dynamically calculated based on the regional illumination to adapt to uneven illumination or reflection interference. Moreover, an illumination normalization layer is added before the adaptive binarization, and Contrast Limited Adaptive Histogram Equalization (CLAHE) is used for processing to enhance the local contrast without amplifying the noise, realizing adaptive illumination compensation.
[0050] Specifically, when performing the preset threshold processing on the second field of view image, the obtained third field of view image is a binary image. The pixel points in the third field of view image that are lower than the preset threshold are marked as target pixel points, and adjacent target pixels will form a target pixel set. Please refer to Figure 4 , Figure 4 which is a schematic diagram of a third field of view image provided by an embodiment of the present application. As shown in Figure 4 , taking the black and white keys as an example, all the continuous and adjacent "black" pixel points (i.e., the regions lower than the preset threshold) in the third field of view image form a target pixel set, and each black key will respectively correspond to a target pixel set, enabling the subsequent steps to more accurately identify the target pixel sets as independent target sets while eliminating small interferences.
[0051] It should be understood that the previously retained key contours are further clarified by threshold segmentation. For example, it can sharpen the slight edge blur caused by the smoothing in step S302, suppress the local gray-scale interference caused by uneven illumination, and initially close the microscopic fracture regions inside the black keys due to texture or uneven illumination. At the same time, the contrast between the keys and the surrounding environment is enhanced. For example, the difference between the black keys and the white keys in the piano is further differentiated, providing a high-confidence binary input for subsequent morphological operations.
[0052] In some embodiments, before step S303, it further includes: increasing the chromaticity-weighted grayscale of the second field of view image, that is, enhancing the difference weight between the green and blue channels. It should be understood that for the black-and-white color tendency, at the junction of the black and white keys, the average chromaticity difference between green and blue is significantly higher than that between red and green, and between red and blue. And under strong light, the chromaticity difference between the black and white keys may be more significant than the brightness difference. Therefore, enhancing the green and blue channels can maximize the chromaticity difference, and more effective features can be retained in subsequent processing.
[0053] S304, perform a second filtering operation on the target pixel set in the third field of view image based on a preset structural element to suppress the image noise of the second form, and obtain a fifth field of view image.
[0054] Among them, the preset structural element is a preset set of pixel points used to detect and correct the shape defects of the target pixel set. Its size and shape need to adapt to the physical characteristics of a type of key, such as the width of a type of key. Correspondingly, the second filtering operation is a morphological operation based on the preset structural element, and the main operations include dilation, erosion, opening operation, and closing operation, etc., to achieve functions such as noise removal, shape feature enhancement or weakening.
[0055] Specifically, based on the structured processing of morphological operations, the image noise of the second form in the third field of view image is repaired by the preset structural element. Among them, the image noise of the second form refers to the non-discrete and continuous interference existing in the image, which is the structural noise caused by shooting conditions or physical structures, such as the fine gaps formed inside the keys due to wood texture or uneven illumination, or the linear interference generated by key reflection, the key gaps of non-a-type keys, or dust accumulation.
[0056] It should be understood that in the second stage, the microscopic fractures inside the black keys caused by the shooting angle are filled by pixel aggregation to reconstruct a complete closed region, and then refined boundary trimming is implemented to synchronously eliminate structural noise and background interference, realizing the synchronous optimization of repairing effective features and stripping invalid features, and avoiding the generation of multi-segmented fractured edges in the unclosed contours during subsequent extraction, which may lead to confusion in the contour hierarchical relationship.
[0057] Please refer to Figure 5 , Figure 5is a schematic diagram of a fifth field of view image provided in an embodiment of the present application, such as Figure 5 As shown in the figure, after the second filtering operation, the fifth field of view image is processed, and the burrs on the edge of a type of piano key (such as the bottom edge of the black piano key), the key seam of the white piano key, and multiple linear interferences in the background are eliminated, thereby achieving refined boundary trimming. At the same time, the interior of the black key is presented as a complete closed continuous state, such as Figure 4 As shown in FIG. 1 , the rightmost black key is disturbed by the reflective area, resulting in multiple breaks of the black key after the preset threshold processing, such as Figure 5 As shown, the gap caused by reflection on the rightmost black key in the fifth field of view image has been corrected to a completely closed continuous state, thereby avoiding the generation of multiple broken edges of the unclosed contour in the subsequent extraction process, which in turn leads to confusion in the contour hierarchy relationship. At the same time, the original geometric features of the keys are retained, that is, the rightmost black key is still in an incomplete form.
[0058] It should be understood that in order to ensure that the key details are fully preserved, the interference filtering strength of the second stage is still strictly limited, e.g. Figure 5 Many black keys are incomplete and the background is still disturbed. This is because the second-stage interference filtering uses conservative parameter settings to avoid excessive interference and damage to the original geometric features of the keys, providing high-integrity and authentic contour data for subsequent perspective correction and physical verification. Furthermore, the first and second stages are both preliminary processing, and fine-tuning is performed on the basis of highly respecting the original image data to eliminate specific image noise.
[0059] In some embodiments, the second form of image noise includes burr noise and linear noise; the S304 includes: based on the first structural element, expanding the boundaries of each target pixel set in the third field of view image to fill the internal holes or breaks of the target pixel set to obtain a fourth field of view image; based on the second structural element, shrinking the boundaries of each target pixel set in the fourth field of view image to eliminate the burr noise of each target pixel set and the linear noise in the fourth field of view image to obtain the fifth field of view image.
[0060] Specifically, a first structural element is used to perform a boundary expansion operation on the target pixel set in the third field of view image. Through morphological dilation, the gaps between adjacent pixels are filled to repair the holes or breaks in the target pixel set caused by threshold processing or uneven illumination, such as the local missing areas in the black key texture, thereby obtaining a fourth field of view image, in which the continuity of the target pixel set is enhanced, but the edges may have slight redundancy due to the expansion. Further, a second structural element is used to perform a boundary contraction operation on the fourth field of view image. Through morphological erosion, the burr noise at the edges of the target pixel set and the residual linear noise in the image are eliminated. In the finally generated fifth field of view image, the contour of the target pixel set is smoother and more compact, while retaining the overall structural characteristics of the keys. Among them, the burr noise refers to the small and irregular protrusions caused by preset threshold processing or uneven illumination, such as the serrated defects at the edges of the keys. The linear noise refers to the slender stripe-shaped noise formed in the image due to light interference or sensor error, as well as other line interferences in the background.
[0061] It should be understood that the preset structural elements include the first structural element and the second structural element. In order to ensure that the key details are fully retained, more conservative first and second structural elements are selected. For example, they are specifically designed according to the physical characteristics and noise characteristics of a type of keys, and the shape and size of their pixel point sets are strictly restricted. And the first structural element and the second structural element can be set to different structures according to their different functions. The specific values can be flexibly set according to the actual situation and are not limited here.
[0062] S305, based on the second preset model, identify several of the type of keys from the fifth field of view image.
[0063] Specifically, by inputting the purified target pixel set in the fifth field of view image (such as the continuous and unbroken key area), the second preset model combines the pre-trained feature library (such as key shape, size, arrangement rule) or rule algorithm (such as contour matching, ratio analysis) to accurately locate and classify the type of keys that meet the preset features. For example, if the target is a black key, the model will identify the narrow and long dark areas arranged regularly in the image. It should be understood that since the previous processing has significantly purified the image, the second preset model only needs to match through simple rules or pre-trained features to identify the type of keys from the continuous target pixel set. Therefore, the second preset model is a pre-configured recognition algorithm or a trained classification model, and its feature library or rules are constructed based on the prior information such as the geometry and arrangement of a type of keys. Existing algorithms or contour matching models can also be used, which are not limited here. Among them, the recognition result of the type of keys can be the key contour and / or the position coordinates of the keys in the image.
[0064] In some embodiments, since some noise or interference information is allowed to be transmitted to subsequent stages in the foregoing processing steps, there may be problems such as local defects and unsmooth contours in the extracted contours of a type of key. Please refer to Figure 6 , Figure 6 which is a schematic diagram of the extraction of the contour of a type of key provided by an embodiment of the present application. As shown in Figure 6 , the red line is the contour of a type of key extracted. Due to the significant effect of the previous denoising stage, the second preset model accurately and without omission extracts each type of key (such as the black keys in the figure), but the contours of multiple black keys are not smooth; there are local parts of some keys (such as the rightmost black key) that are not correctly recognized. Based on this, an embodiment of the present application proposes a contour correction mechanism based on perspective deformation.
[0065] In some embodiments, the method further includes: based on preset camera parameters, determining the perspective deformation parameters of each type of key according to the position coordinates of the type of key in the fifth field of view image; determining the target shape parameters of each type of key according to the reference shape parameters of the type of key and the perspective deformation parameters, and adjusting the contour of the corresponding type of key according to the target shape parameters of each type of key.
[0066] Specifically, based on the preset camera parameters and the position coordinates of a type of key in the fifth field of view image, the perspective deformation parameters (such as the contraction ratio) of each type of key are determined, and the reference shape parameters of the key are adjusted according to the perspective deformation parameters to simulate the perspective effect, so that the reference shape parameters are deformed due to perspective, such as edge stretching or shortening. Finally, the key contour is corrected according to the target shape parameters of each key to ensure that the calibrated key contour restores the deformation error caused by the shooting angle, thereby improving the accuracy of subsequent positioning.
[0067] It should be noted that for the same preset camera parameters, the perspective deformation parameters at each position coordinate are determined. Based on the preset camera parameters and the position coordinates of a type of key, the corresponding perspective deformation parameters can be quickly obtained. Among them, the perspective deformation parameters at each position coordinate in the image are determined based on prior data and are not limited herein. Further, the target shape parameters can be derived through affine transformation or projective transformation to obtain the target shape parameters of each key.
[0068] In some embodiments, the reference shape parameters can be the geometric feature parameters of the key in the image without perspective distortion, such as width, length, length-width ratio, etc. The reference shape parameters can be obtained in advance, or the shape parameters of the key with the lowest degree of perspective distortion in the image can be determined as the reference shape parameters.
[0069] In some embodiments, the method further includes: obtaining at least one proximal first-type key with a perspective deformation parameter lower than a first deformation threshold, and determining a reference shape parameter of the first-type keys based on the shape parameters of the at least one proximal first-type key; obtaining at least one distal first-type key with a perspective deformation parameter higher than a second deformation threshold, where the first deformation threshold is less than or equal to the second deformation threshold; adjusting the reference standard shape parameter of the first-type keys based on the perspective deformation parameter of each distal first-type key to determine the target shape parameter of each distal first-type key; and adjusting the contour of the corresponding distal first-type key according to the target shape parameter of each distal first-type key.
[0070] Specifically, proximal first-type keys with perspective deformation parameters lower than the first deformation threshold are screened out (such as keys directly in front of the camera with almost no perspective distortion), and their shapes are close to the reference shape in the standard view. By statistically analyzing the shape parameters of at least one proximal key, the reference shape parameter of the first-type keys is calculated and updated for subsequent adjustment. For distal first-type keys with perspective deformation parameters higher than the second deformation threshold (such as keys at the edge of the image with obvious deformation due to inclination), based on their specific perspective deformation parameters, the reference shape parameter is dynamically adjusted through perspective transformation. For example, if the width of a key is shortened due to inclination at a distance, the reference shape parameter is adjusted to synchronously shorten the width to generate the target shape parameter. The first deformation threshold is used to identify proximal first-type keys less affected by perspective deformation. Correspondingly, the second deformation threshold is used to identify distal keys affected by perspective deformation, and the specific value can be flexibly set and is not limited here.
[0071] According to the target shape parameter of each distal key, the contour shape in the image is adjusted (such as stretching or scaling the pixel region in a specific direction) to make it closer to the perspective shape. Thus, through sub-region processing of proximal and distal parts and dynamic parameter correction, the key deformation caused by the camera view difference is adapted in contour calibration to accurately locate keys at various positions in the image.
[0072] In some embodiments, when there is no perspective, the scaling ratio in the perspective deformation parameter is 1.0 times, and the target shape parameter is the reference shape parameter. As the first-type keys gradually move away from the optical axis in the image, the scaling ratio shows a linear change during the process and gradually reaches 0.5 times, where the target shape parameter is the reference shape parameter.
[0073] In some embodiments, the contours of several first-type keys are obtained, and based on the width of the contours of the first-type keys, keys with a width greater than a width threshold are obtained, and the key with the narrowest width is selected from them. The shape parameter of the key with the narrowest width is used as the reference shape parameter. The width threshold can be flexibly set and is not limited here.
[0074] In some embodiments, a limit value is set for the target form parameter. When the target form parameter obtained by adjusting the reference form parameter is less than the limit value, the limit value is taken as the target form parameter; and / or, the number of anomalies where the target form parameter obtained by adjusting the reference form parameter is less than the limit value is counted. If the number of anomalies is greater than a preset number of anomalies, a proximal type of key is reselected to update the reference form parameter. For example, when key A is selected as the proximal type of key and the target form parameter obtained by adjusting the reference form parameter of key A is less than the limit value, the limit value is taken as the target form parameter. When the target form parameter is less than the limit value for multiple consecutive times, key B is reselected as the proximal type of key.
[0075] It should be understood that in the third stage, the form data of the proximal keys is used as the reference parameter in combination, and the distal keys are processed through perspective deformation parameter grading, which not only maintains the reliability of parameter calibration but also avoids parameter distortion caused by perspective projection. Even when there are differences in musical instrument production processes, changes in musical instrument models, or changes in shooting angles, a high recognition rate can still be maintained, and it has stronger generalization ability.
[0076] In some embodiments, please refer to Figure 7 , Figure 7 is another schematic diagram of the extraction of a type of key wheel provided by an embodiment of the present application. As Figure 7 shown, the red lines are the outlines of the extracted type of keys. The second preset model accurately and without omission extracts each type of key. At the same time, due to strong background interference, the second preset model misidentifies non-key areas as keys, such as the striped carpet and long rectangular frame in the figure. At this time, misjudgments can be further filtered through the position verification or arrangement rule analysis in the fourth stage to ensure the accuracy of the final result. Therefore, the extraction of the type of keys in step S305 is also a kind of filtering of image interference. At this time, the filtering intensity is still strictly limited. At this time, a high-quality fifth field of view image is used to achieve efficient preliminary screening, rather than pursuing absolute accuracy, and a certain recognition error is tolerated.
[0077] In some embodiments, the method further includes: obtaining the actual form parameters and position coordinate parameters of each type of key; based on preset key layout parameters, screening a plurality of the type of keys according to the actual form parameters and position coordinate parameters of the type of keys to determine a plurality of target type of keys; and outputting the position coordinates of the plurality of target type of keys in the first field of view image.
[0078] Among them, the preset key layout parameters are a set of standardized physical parameters of a type of keys in the image, which are used to verify the geometric and arrangement rationality of the recognition results. Exemplarily, the preset key layout parameters include at least one of the target length, target width, target length-width ratio, target area, target key spacing, and target number of keys of the type of keys. For example, the total number of black keys in a standard piano is 36.
[0079] Specifically, for each type of key obtained by recognition or each type of key whose key contour is adaptively corrected based on perspective deformation, the actual morphological parameters and position coordinate parameters presented by it are obtained, and multi-level verification is performed based on the preset key layout parameters. First, in combination with the preset key layout parameters (target length, target width, target length-width ratio, target area, etc.), the morphological verification of a single key is carried out, strictly screening the areas that conform to the physical characteristics of the key, removing the type of keys whose size, length-width ratio or area deviate significantly, and excluding the interfering objects with abnormal sizes; second, through the key spacing, quantity and arrangement mode, the overall layout verification of the keyboard is carried out, retaining the type of keys that conform to the preset arrangement rules and the number of keys, ensuring that the recognition result conforms to the standard structure of the musical instrument. Thus, the target type of keys that meet the standard are accurately retained, and the non-key areas misrecognized in the background are removed, and the accurate position coordinates of the target type of keys in the original first field of view image are output, providing reliable data for subsequent positioning or interaction operations.
[0080] In the embodiment of the present application, the interference filtering intensity of step S302 is strictly limited, and only discrete noise is preliminarily purified. While retaining the clarity of the main structure of the keys, it also retains complex environmental interferences such as local breaks caused by the shooting angle, uneven illumination or the texture of the black keys. And it is strictly limited to complete the removal of discrete isolated noise before the threshold processing (step S303), ensuring that the edge blurring that may be caused by the smoothing strategy is clearly restored during the threshold processing, ensuring the geometric integrity of the keys, and further providing a binary input with high confidence for subsequent morphological operations. Further, step S304 repairs the breaks of the keys or edges that appear after the threshold processing while removing continuous noise, ensuring that the key contour is closed, and avoiding the generation of multi-segment broken edges by the unclosed contour during the subsequent extraction process, which may lead to confusion in the contour hierarchical relationship. Since the interference filtering in the second stage uses conservative parameter settings, the processed keys may be incomplete, and the background may still have interference.
[0081] In step S305, the extraction of the type of keys uses the high-quality fifth field of view image to achieve efficient preliminary screening, which can filter most of the background interference. However, the filtering intensity at this time is still strictly limited, and it does not pursue absolute accuracy, but tolerates a certain recognition error, such as contour defects of the type of keys and misjudgment of non-key areas. Please refer to Figure 8 , Figure 8This is a schematic diagram of the adjusted contour of a certain type of key after the implementation of this application. As Figure 8 shown, for these recognition errors, they are eliminated by adaptively correcting the key contour based on perspective deformation and performing hard physical verification based on key layout parameters. Figure 8 In the figure, the green line is the adjusted contour of a certain type of key. While accurately fitting the certain type of key, the contour presents a smooth and regular form. Also, to avoid premature pruning that may cause the effective contour to be mistakenly deleted due to local deformation, the physical verification is carried out at the end of the process.
[0082] Therefore, each step in the implementation of this application has high timing and coupling. Through a phased and progressive interference filtering mechanism, even if some noise or fuzzy information is tolerated and transmitted to subsequent stages, it will be gradually recognized and corrected in subsequent steps. Finally, high-precision interference elimination is achieved while retaining key details, effectively improving the recognition accuracy and robustness of keys in complex scenarios.
[0083] Please refer to Figure 9 , Figure 9 This is a schematic flowchart of a key recognition method based on key slots provided by the implementation of this application. As Figure 9 shown, the implementation of this application provides a key recognition method based on key slots; the method includes S401 to S404.
[0084] S401, obtain a first field of view image, where the first field of view image includes images corresponding to several certain types of keys and other types of keys.
[0085] Among them, the first field of view image refers to the original image directly collected by the image acquisition module. In this image, both certain types of keys and other types of keys are presented simultaneously, and it is required to contain sufficient details to support subsequent positioning and segmentation operations, such as clear key slot boundaries or key contours, as the input basis for subsequent key recognition.
[0086] S402, based on the preset positional relationship between the certain types of keys and the other types of keys, and according to the key image area where the certain types of keys are located, intercept several first partial images of the areas where the other types of keys are located from the first field of view image.
[0087] Among them, the preset positional relationship refers to the fixed spatial arrangement rule between certain types of keys and other types of keys. Taking a piano as an example, the black keys are usually located in the upper gaps between adjacent white keys, and there is a fixed size ratio and specific interval arrangement between them, such as two black keys sandwiched between three white keys.
[0088] Specifically, based on the key image area where a known type of keys is located, according to the preset positional relationship, a first partial image containing only the second type of keys is intercepted from the first field of view image. Thus, the processing range is quickly narrowed through spatial correlation, converging the global processing to the local, reducing the computational amount and excluding the interference of irrelevant backgrounds, while retaining the complete details of the target area, providing a high-precision input for subsequent key gap recognition.
[0089] Please refer to Figure 10 , Figure 10 which is a schematic diagram of intercepting the first partial image provided by an embodiment of the present application. As Figure 10 shown, the musical instrument is a piano, the image acquisition module is vertically arranged on the key plane and is located directly above the keyboard. The first field of view image includes several first type of keys 100 (such as the black keys in the figure) and several second type of keys 200 (such as the white keys in the figure). According to the preset positional relationship, a rectangular area is intercepted below and on both sides of the first type of keys 100, such as the red dotted line area, covering the area where the possible second type of keys 200 may exist.
[0090] Please refer to Figure 11 , Figure 11 which is a schematic diagram of recognizing keys based on key gaps provided by an embodiment of the present application. As Figure 11 shown in 40a in the figure, the region of interest (ROI) of the white keys, that is, the first partial image, is intercepted, which includes the partial images of several second type of keys, quickly excluding the interference of irrelevant backgrounds. It should be understood that there are significant differences in color, shape or texture between the first type of keys and the second type of keys, such as the contrast between dark and light colors. Thus, the first partial image can be quickly intercepted through the first type of keys.
[0091] In some embodiments, the regions that meet the size and shape of the first type of keys can be screened by combining shape features (such as rectangles, aspect ratios); the template image of the first type of keys (such as local texture or contour) can also be pre-stored and slid and matched in the target image to find the region with the highest similarity; a detector for the first type of keys can also be trained through a large number of labeled data to output the position and category. The features of the first type of keys are more obvious than those of the second type of keys, and its recognition accuracy is higher. The recognition method of the first type of keys can refer to the related technology and will not be limited here.
[0092] In some embodiments, through a key recognition method provided by any embodiment of the present application, the first type of keys and the position coordinates of the first type of keys in the first field of view image can be recognized, and then the key image area where the first type of keys is located can be determined.
[0093] In some embodiments, in a common musical instrument keyboard, the keys are divided into two categories: black keys and white keys. The black keys are regarded as one type of keys, and the white keys are regarded as another type of keys. Preferably, a first partial image is intercepted in a region based on the coordinates of the black keys, and then white key recognition is performed on the basis of the first partial image. Compared with the white keys, the features of the black keys are more obvious. Performing global processing can avoid introducing a large amount of noise and improve the effectiveness of calculation. At the same time, it is difficult to distinguish the projection of the black keys from the edges of the white keys in the global image. Local processing of the white keys can reduce the false detection rate.
[0094] In some embodiments, a preset positional relationship is determined based on the prior knowledge of a standard musical instrument. Exemplarily, the preset positional relationship between the one type of keys and the other type of keys is the shape ratio or area ratio between the one type of keys and the other type of keys. Among them, the shape ratio is the relative relationship between the two types of keys in terms of geometric shape, such as the ratio of length, width or height. For example, in a piano, the length of the white key is generally three-halves of the length of the black key. The area ratio is the area ratio between the two types of keys in the image or the actual physical structure.
[0095] In some embodiments, after S402, the method further includes: performing a preset threshold processing on the first partial image and performing a first filtering operation to suppress image noise of a first form; or, performing a first filtering operation on the first partial image to suppress image noise of a first form and performing a preset threshold processing.
[0096] Specifically, the preset threshold processing is a binary processing corresponding to a global or adaptive threshold, which separates the keys in the first partial image from other backgrounds (such as key gaps); the first filtering operation is to suppress the image noise of the first form through a multi-directional smoothing strategy. Preferably, the multi-directional smoothing strategy is Gaussian filtering. The specific implementation manners can refer to the foregoing embodiments and will not be elaborated herein. Moreover, the specific parameter settings of the preset threshold processing and the first filtering operation can be the same as or different from those of the foregoing embodiments, and are not limited herein. And, since the background interference in the first partial image is extremely low, and the key gap features are clear and not easy to be blurred, the two processes can be flexibly combined in sequence and are not limited herein. The processed first partial image is as Figure 11 shown in 40b.
[0097] S403, identifying a plurality of initial key gaps between adjacent other type of keys in the first partial image; and performing a morphological operation on the plurality of initial key gaps in the first partial image to obtain a target partial image.
[0098] Specifically, the initial key gaps between adjacent other type of keys are identified from the first partial image, such as low-brightness regions or boundary lines. As Figure 11Among them, 40c is the recognition result of the initial key gap. The lines corresponding to the white pixels are the initial key gaps. Further, morphological operations are performed on these initial key gaps to eliminate structural noise and standardize the gap morphology. For example, Figure 11 Among them, 40d is the target local image output after morphological operation processing. The key gaps have been corrected into continuous strips that conform to physical laws (such as uniform width and perpendicular direction), while retaining the arrangement relationship between the key gaps and the keys, providing a high-confidence input for the key area extraction of S404.
[0099] It should be understood that the key gap (also known as the key clearance) is the natural boundary area between adjacent second-class keys in the first local image, manifested as a narrow strip with low brightness or high contrast. The recognition method includes local feature analysis based on color contrast (such as the brightness difference between the key gap and the key) or edge detection. And, for example, Figure 11 As shown in 40c, the image area corresponding to the initial key gap is very narrow. When performing morphological processing, the corresponding structural element needs to be strictly defined to avoid overfilling the gap or destroying the original gap information.
[0100] In some embodiments, S403 includes: identifying the contours of several second-class keys from the first local image according to the color information of several pixels in the first local image; identifying several candidate key gaps between the second-class keys from the first local image; obtaining at least one adjacent second-class key adjacent to the candidate key gap; and determining the gap inclination threshold of the adjacent second-class keys based on a preset angle constraint according to the longitudinal contours of the adjacent second-class keys. When the inclination angle of the candidate key gap meets the gap inclination threshold, the corresponding candidate key gap is used as the initial key gap.
[0101] Specifically, based on the color characteristics of the pixels in the first local image, such as gray value or color saturation, narrow lines with sudden brightness changes are detected as candidate key gaps. These candidate key gaps not only include the real key gaps but may also include pseudo-gaps caused by reflection, stains, or perspective deformation, such as the oblique bright lines formed by reflection. Obtain the adjacent second-class keys adjacent to the candidate key gap, such as the second-class keys on the left and / or right sides of the candidate key gap, or the second-class keys within a preset adjacent range of the candidate key gap. The perspective deformation degrees of the candidate key gap and the adjacent second-class keys are related and similar. Analyze the actual arrangement direction of the longitudinal contours of the adjacent second-class keys, calculate the tolerable gap inclination threshold in combination with the preset angle constraint, and then screen several candidate key gaps, retaining the initial key gaps whose inclination degrees meet the requirements, while those exceeding the threshold are determined as noise or pseudo-gaps and excluded.
[0102] Among them, the preset angle constraint is the tolerable angle deviation set in advance, and the specific value can be flexibly set according to actual needs, such as 15°, which is not limited here.
[0103] Among them, the longitudinal contour refers to the edge contour of the second-class keys extending along their long axis direction, such as the straight edges on both sides of the white keys, which is used to determine the relative angle between the main extension direction of the keys and the image coordinate system. It should be understood that since the longitudinal contour of the second-class keys may be tilted as a whole due to the pitch angle of the camera or the longitudinal contour may be tilted due to perspective, therefore, dynamically calculating the gap tilt threshold according to the actual longitudinal contour direction can conform to the actual key gap direction in the image. For example, if the actual longitudinal contour direction is tilted 5° to the right and the preset angle constraint is 15°, then the gap tilt threshold is tilted 10° to the left to 20° to the right, and the candidate key gaps with tilt angles within this range are used as the initial key gaps.
[0104] Therefore, by combining the key contour direction and the dynamic angle constraint, it is ensured that the recognized key gaps conform to the actual tilt angle of the keys in the image, and only the key gaps that conform to the perspective law of the specific image position are retained, effectively excluding interference information, and providing an accurate data basis for subsequent key gap repair and key positioning.
[0105] In some embodiments, based on the geometric characteristics of the parallel arrangement of the second-class keys, a vertical edge enhancement strategy is adopted. Through a customized horizontal direction gradient filter, the vertical gap features between the second-class keys are preferentially captured, while the oblique interference edges are suppressed, the image gradient distribution is dynamically analyzed, and high and low thresholds are automatically set to ensure the balance between weak edge retention and noise suppression. At the same time, the direction of the detected edges (i.e., candidate key gaps) is screened, and only the effective edges with a deviation from the arrangement direction of the second-class keys less than 15° are retained, and approximately vertical lines are screened. Thus, oblique interferences such as the projection and reflection of the first-class keys are eliminated through angle constraints, ensuring the physical consistency of the edges.
[0106] In some embodiments, the morphological operation includes the steps of: S4031, obtaining at least one target second-class key adjacent to the initial key gap to be processed, and updating the setting parameters of the structural element based on the perspective deformation parameters of the target second-class key.
[0107] Specifically, for each initial key gap to be processed, or for multiple initial key gaps within a preset common range, set parameters of a specific structural element are used. By analyzing at least one target second-class key adjacent to the initial key gap, a perspective deformation parameter of the target second-class key is obtained, and the set parameters of the structural element used in the morphological operation are dynamically calculated and updated to match the actual deformation characteristics of the key gap. It should be understood that due to perspective, the keys are in a tilted or longitudinally contracted state, resulting in a deviation between the direction of the adjacent key gap and the standard horizontal direction. At this time, according to the specific perspective deformation of each key, the structural element is adaptively adjusted to differentially process different key gaps.
[0108] In some embodiments, the difference in perspective deformation parameters of multiple initial key gaps within a preset common range is less than a preset deformation difference, and a set of parameters of a structural element can be shared. The specific range division can be determined in advance and will not be limited here.
[0109] In some embodiments, the target second-class key can be the second-class key on the left and / or right side of the initial key gap. When multiple initial key gaps share a set of parameters of a structural element, the second-class key on the left and / or right side of any one of the initial key gaps can be selected as the target second-class key.
[0110] For example, the long axis direction of the structural element needs to be aligned with the trend of the adjacent target second-class key so that the processed key gap matches the shape of the adjacent target second-class key. The size parameters of the structural element can be adjusted according to the deformation degree of the key, such as the length of the structural element shortens with perspective contraction.
[0111] It should be understood that based on the dynamic coupling mechanism between the perspective deformation parameter and the structural element parameter, the morphological operation can accurately repair key gap breaks or eliminate noise, avoiding processing deviations caused by fixed structural elements, thereby generating a target local image that more conforms to the key deformation law in the image.
[0112] In some embodiments, the position coordinates of the target second-class key in the first field of view image are obtained, and based on the preset camera parameters and the position coordinates of the target second-class key, the perspective deformation parameter of the target second-class key is determined. For example, the position coordinates of the target second-class key in the first field of view image are obtained, and its perspective deformation parameter is calculated in combination with the preset camera parameters. Another example is that for the same preset camera parameters, the perspective deformation parameter at each position coordinate is determined. Based on the preset camera parameters and the position coordinates of the target second-class key, the corresponding perspective deformation parameter can be quickly obtained, where the perspective deformation parameter at each position coordinate in the image is determined based on prior data and will not be limited here.
[0113] In some embodiments, the perspective deformation parameters include a contraction ratio and / or a contraction direction. The contraction ratio refers to the specific numerical ratio of the deformation of an object (such as a type of key, a target type-two key) in the image due to perspective deformation. For example, the contraction ratio in the length direction and the width direction compared to the non-perspective image. The contraction direction refers to the direction in which the object (such as a type of key, a target type-two key) is shortened in the image due to perspective deformation. For example, since the distal key is far from the camera, its longitudinal direction appears shortened in the image, that is, the contraction direction is the depth direction perpendicular to the line of sight.
[0114] In some embodiments, the setting parameters of the structural element include a major axis direction, a major axis length, and a minor axis length. The method further includes: obtaining the perspective deformation parameters of the target type-two key, where the perspective deformation parameters include a contraction ratio and / or a contraction direction; determining the major axis direction of the structural element based on the contraction direction; and / or, determining the target morphological parameters of the target type-two key according to the standard morphological parameters of the type-two key and the contraction ratio; determining the major axis length and the minor axis length of the structural element according to the target morphological parameters of the target type-two key.
[0115] Specifically, the major axis direction of the structural element is dynamically adjusted according to the contraction direction of the target type-two key. If the key is longitudinally contracted due to perspective, its contraction direction is the inclination direction of the major axis of the key, which is also the inclination direction of the key gap. The major axis direction of the corresponding structural element needs to be strictly aligned with this contraction direction to ensure that the morphological operation is consistent with the actual trend of the key gap in the image. For example, the distal key shows a shortening trend in the image, and the key is inclined to the upper right, then the major axis of the structural element rotates along the same direction.
[0116] Specifically, obtain the standard morphological parameters of the type-two key, such as the original length and width when no perspective deformation occurs, and calculate its target morphological parameters, such as the actual length and width after being shortened due to perspective. Further, the major axis length of the structural element is dynamically scaled and determined according to the actual length in the target morphological parameters, and / or, obtain the ratio setting between the type-two key and the key gap width, determine the expected width of the key gap according to the actual width in the target morphological parameters and the ratio setting, and the minor axis length of the structural element is dynamically scaled and determined according to the expected width of the key gap, ensuring that the morphology of the structural element highly matches the physical characteristics of the actual key gap. Thus, by combining the perspective deformation parameters with the standard morphological parameters, the size parameters of the structural element are accurately adapted to the key morphology under the current perspective, thereby improving the accuracy of key gap repair and noise suppression.
[0117] In some embodiments, the morphological operation further includes the step of: S4032, performing a morphological operation on the initial key gap to be processed based on the updated structural element.
[0118] Specifically, using the updated structural element, morphological operations are performed on multiple initial key - to - key gaps within one or a preset common range to optimize the gap morphology. Through dynamically adapted morphological processing, the structural element can accurately fit the actual key - gap trend and deformation law, retaining effective features while suppressing noise. As shown in 40d of FIG. 11, the morphological operation thickens, expands, and dilates the white - key edge (i.e., the initial key - to - key gap), and finally generates a target local image that only contains complete, continuous, and physically regular key - to - key gaps, improving the robustness and accuracy of subsequent key positioning.
[0119] In some embodiments, the method further includes: expanding the boundary of the initial key - to - key gap based on the updated structural element to fill internal holes or breaks in the initial key - to - key gap, obtaining a second local image; shrinking the boundary of the initial key - to - key gap in the second local image based on the updated structural element to eliminate burr noise and linear noise in the initial key - to - key gap, obtaining the target local image.
[0120] Specifically, the key - gap features are optimized through two - stage morphological operations. First, based on the updated structural element, an expansion operation is performed on the boundary of the initial key - to - key gap, such as morphological dilation. By sliding the structural element, the internal holes or broken areas in the key - gap are filled to generate a second local image, ensuring the continuity of the key - gap. Second, the structural element is used again to perform a contraction operation on the key - gap boundary of the second local image, such as morphological erosion, to eliminate edge expansion or pseudo - key - gaps caused by excessive expansion, and finally form a target local image that only retains complete, smooth, and physically regular key - to - key gaps.
[0121] In some embodiments, the initial key - to - key gap is traversed based on the updated structural element to expand the pixel set that matches the structural element and / or contract the pixel set that does not match the structural element. For the pixel set that matches the morphology of the structural element, such as the pixel set with a direction - adapted key - gap tilt angle and a size - matched perspective scaling ratio, an expansion operation is performed to repair key - gap breaks and key - key adhesions caused by image noise or perspective tilt; while for the pixel set that does not match the structural element, such as the pixel set with a direction - inadapted key - gap tilt angle, a contraction operation is performed to eliminate redundant or deviated - from - expected morphological regions, such as edge burrs or pseudo - gaps formed by background interference.
[0122] In some embodiments, the structural element is in an elliptical shape and, as the long - axis direction changes, presents an inclined elliptical morphology, that is, an inclined elliptical kernel. Exemplarily, the inclined elliptical kernel is an asymmetric and direction - adjustable morphological structural element. It is in an elliptical shape, and the lengths of the long axis and the short axis are dynamically calculated according to the actual physical dimensions of the second - type keys and the camera parameters. The long - axis direction is consistent with the perspective contraction direction of the second - type keys, enhancing the continuity of the key - gap. The kernel parameters change with the image position to match the deformation degree of different regions.
[0123] In some embodiments, according to the actual physical size of the second - type keys (about 23 mm wide) and the preset camera parameters (such as focal length, tilt angle), a mathematical model of the kernel size is constructed. In the image edge deformation region, an inclined elliptical kernel is used for gap connection, and the long - axis direction is consistent with the perspective contraction direction of the white keys to achieve geometric compensation.
[0124] It should be understood that the closer the keys are to both sides of the image, the smaller the key gaps are. And in the case of relatively high light intensity, it may cause two white keys to stick together. The essential process of morphological analysis is dilation and erosion. In the embodiments of the present application, the kernel that determines the degree of dilation and erosion is dynamically updated, avoiding using the same kernel for processing, thereby avoiding the kernel being too large and shrinking the judgment area of the white keys; and the kernel being too small to separate the stuck - together white keys.
[0125] It should be understood that a coupling mechanism of dynamic structural elements and perspective deformation parameters is introduced. According to perspective deformation parameters such as the tilt angle and contraction direction of the keys, the morphological parameters of the structural elements (such as long - axis direction, long - and short - axis lengths) are adjusted in real time to accurately match the key - gap trend and morphological changes, effectively improving the accuracy and anti - interference ability of key recognition. For example, if the key - gap trend of the target second - type keys deviates from the vertical direction due to the tilt of the camera, the long - axis direction of the structural element is adjusted to be consistent with the tilt angle of the key gap, ensuring that the morphological operation is carried out along the deformation direction of the key gap to accurately repair the broken or peeled - off noise; if the key - gap width is scaled due to perspective projection of the keys, the size of the structural element is reduced according to the scaling ratio to avoid over - filling or erosion, enabling the structural element to adapt to the key - gap deformation under different shooting conditions and taking into account the geometric feature consistency of the proximal and distal key gaps, thereby improving the robustness of the morphological operation.
[0126] S404, identify a plurality of the second - type keys from the target local image based on the third preset model.
[0127] Specifically, by inputting the complete, continuous and noise - free key - to - key gaps in the target local image, clear key - boundary features are provided for the third preset model. The third preset model combines a pre - trained feature library (such as key shape, size, arrangement rule) or a rule - based algorithm (such as contour matching, ratio analysis) to accurately locate and classify the second - type keys that meet the preset features. For example, if the target is a white key, a wide and continuous bright - color area is extracted. As Figure 11 As shown in 40e, the key - image regions of a plurality of second - type keys extracted from the target local image are compared and displayed in the first local image, where the red lines are the contours of the extracted second - type keys, and it can be seen that the third preset model accurately extracts each second - type key without omission.
[0128] It should be understood that since the key slots are repaired for fractures, burrs are eliminated, and the perspective deformation of the keys is strictly matched through dynamic structural elements, the highly regular boundary contour of the key slots means that the boundary contour of the keys is highly regular. The third preset model can identify the second-class keys from the target local image only through simple rules or pre-trained feature matching. Therefore, the third preset model is a pre-configured recognition algorithm or a trained classification model, and its feature library or rules are constructed based on the prior information such as the geometry and arrangement of the first-class keys. Existing algorithms or contour matching models can also be used, which are not limited herein.
[0129] In some embodiments, the method further includes: obtaining at least one proximal key slot with a perspective deformation parameter lower than the first deformation threshold, and determining the reference morphological parameter of the key slot based on the morphological parameters of the at least one proximal key slot; obtaining at least one distal key slot with a perspective deformation parameter higher than the second deformation threshold, where the first angle threshold is less than or equal to the second angle threshold; adjusting the reference morphological parameter of the key slot based on the perspective deformation parameter of each distal key slot to determine the target morphological parameter of each distal key slot; and adjusting the contour of the corresponding distal key slot according to the target morphological parameter of each distal key slot.
[0130] Specifically, proximal key slots with perspective deformation parameters lower than the first deformation threshold are selected (such as the slots between the second-class keys directly in front of the camera with almost no perspective distortion), and their shapes are close to the reference shape under the standard perspective. By statistically analyzing the morphological parameters of at least one proximal key slot, the reference morphological parameter of the key slot is calculated and updated for subsequent adjustment reference. For distal key slots with perspective deformation parameters higher than the second deformation threshold (such as the slots between the second-class keys with obvious deformation due to inclination at the edge of the screen), based on their specific perspective deformation parameters, the reference morphological parameter is dynamically adjusted through perspective transformation. For example, if the length of the distal key slot is shortened due to perspective, the reference length is scaled proportionally; if the direction is offset, the long axis direction is adjusted to match the inclination angle.
[0131] Based on the adjusted target morphological parameter, the contour of the distal key slot is adjusted, such as Figure 11 as shown in 40f, the green line is the adjusted contour of the second-class keys. While accurately fitting the second-class keys, the contour presents a smooth and regular shape. Thus, it is ensured that its shape is consistent with the reference features of the proximal key slots, improving the geometric regularity and recognition accuracy of the overall key contour, and adapting to the deformation law brought by its own perspective. Furthermore, the subsequently extracted second-class keys can also adapt to the deformation law brought by their own perspectives.
[0132] In the embodiments of the present application, a type of keys with significant features is preferentially recognized. Based on the spatial correlation between keys (such as shape ratio or area ratio), the local image of the second type of keys to be recognized is quickly locked, and the global view is converged to the local image, which not only ensures the positioning accuracy but also greatly reduces the image processing range, reduces the computational complexity, and eliminates background interference. Then, the natural boundary between keys, i.e., the key gap, is selected as the core processing unit, and the reverse normalization strategy is adopted to avoid directly processing large homogeneous areas, further reducing the image processing range to the microscopic structure of the key gaps and reducing the computational complexity.
[0133] In some embodiments, the interception of the first local image is preferably performed before the first filtering operation and edge detection (i.e., initial key gap detection) to avoid unnecessary signal-to-noise ratio degradation. Secondly, morphological operations are preferably performed after edge detection to enhance feature connection and reduce the influence of isolated noise using the edge intensity map, thereby avoiding an increase in the key gap breakage rate. Therefore, an incremental path of edge detection, morphological operations, and contour extraction (i.e., recognition of the second type of keys) is adopted to gradually strengthen the features. Finally, physical verification (such as pitch check, aspect ratio filtering) is performed at the end of the process. At this time, the judgment is based on the complete contour information, which can prevent premature pruning and incorrect deletion of valid contours caused by local deformation or information fragmentation.
[0134] In some embodiments, please refer to Figure 12 , Figure 12 is a schematic diagram of the recognition result of keys provided by the embodiments of the present application. As Figure 12 shown, a key recognition method provided by the embodiments of the present application can be used to recognize a type of keys (such as the black keys in the figure), and the recognition result is shown by the red lines. A key recognition method based on key gaps provided by the embodiments of the present application can be used to recognize the second type of keys (such as the white keys in the figure), and the recognition result is shown by the green lines. Further, the recognition result can be the position coordinates, key contours, and the key image areas where several types of keys and several second types of keys are located.
[0135] In some embodiments, in the step of determining the target morphological parameters of each of the first type of keys according to the reference morphological parameters of the keys and the perspective deformation parameters, and adjusting the contour of the corresponding first type of keys according to the target morphological parameters of each of the first type of keys, based on the position coordinates of the corresponding first type of keys in the image, the top surface position of the first type of keys is determined. Based on the top surface position, the contour of the corresponding first type of keys is adjusted according to the target morphological parameters of each of the first type of keys, so that the contour of the corresponding first type of keys includes the top surface position. It should be understood that the three-dimensional height difference of the black keys causes perspective deformation and edge occlusion effects during imaging, resulting in the offset of their top surfaces, as Figure 6As shown, the black key on the far right exhibits the problem of top surface offset. At this time, based on its position coordinates, its top surface position is predicted to be on the right side of the contour. The contour is adjusted with the right side of the contour as the limiting boundary. That is, if it is necessary to shrink the contour (such as adjusting the contour according to the expected width), the left contour will be preferentially shrunk while retaining the right part of the image, such as Figure 12 As shown, the adjusted contour preferentially retains the right image to ensure accurate selection of the top surface of the key. Thus, reliable data is provided for subsequent positioning or interaction operations, improving the adaptability in complex technique analysis, teaching feedback, and intelligent musical instrument interaction.
[0136] Exemplarily, the above method can be implemented in the form of a computer program that can run on an intelligent musical instrument. The intelligent musical instrument includes keys for directly interacting with the performer and playing music; an image acquisition module for obtaining a first field of view image, where the first field of view image includes images corresponding to several first-class keys or the first field of view image includes images corresponding to several first-class keys and second-class keys. The intelligent musical instrument further includes a memory for storing the computer program; and a processor for executing the computer program and implementing the key recognition method and / or the key recognition method based on key seams provided in any embodiment of the present application when executing the computer program.
[0137] Exemplarily, please refer to Figure 13 , Figure 13 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The above method can be implemented in the form of a computer program that can run on a computer device such as Figure 13 shown. As Figure 13 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions that, when executed, can cause the processor to execute any key recognition method and / or the key recognition method based on key seams. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it can cause the processor to execute any key recognition method and / or the key recognition method based on key seams. The network interface is used for network communication, such as sending assigned tasks, etc.
[0138] It should be understood that the processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0139] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps: S301, obtain a first field of view image, where the first field of view image includes images corresponding to a number of first-type keys; S302, perform a first filtering operation on the pixel points in the first field of view image based on the pixel values in the first field of view image to suppress image noise of a first form and obtain a second field of view image; S303, perform a preset threshold processing on the second field of view image to obtain a third field of view image, where the third field of view image includes a number of target pixel sets, and each target pixel set is composed of a number of consecutive pixel points with pixel values lower than the preset threshold; S304, perform a second filtering operation on the target pixel sets in the third field of view image based on a preset structural element to suppress image noise of a second form and obtain a fifth field of view image; S305, based on a second preset model, identify a number of the first-type keys from the fifth field of view image.
[0140] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps: S401. Obtain a first field of view image, where the first field of view image includes images corresponding to a number of first - type keys and second - type keys; S402. Based on the preset positional relationship between the first - type keys and the second - type keys, and according to the key image regions where the first - type keys are located, intercept a number of first partial images of the regions where the second - type keys are located from the first field of view image; S403. Identify a number of initial key - to - key gaps between adjacent second - type keys in the first partial image; and perform morphological operations on the number of initial key - to - key gaps in the first partial image to obtain a target partial image; S404. Based on a third preset model, identify a number of the second - type keys from the target partial image; where the morphological operations include the steps: S4031. Obtain at least one target second - type key adjacent to the initial key - to - key gap to be processed, and update the setting parameters of the structural element based on the perspective deformation parameters of the target second - type key; S4032. Perform morphological operations on the initial key - to - key gap to be processed based on the updated structural element.
[0141] Exemplarily, the processor is used to run a computer program stored in the memory, and is also used to implement the steps of the key recognition method and / or the key recognition method based on key gaps provided in any embodiment of the present application, which will not be elaborated here.
[0142] In an embodiment of the present application, a computer - readable storage medium is also provided. The computer - readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement the steps of any one of the key recognition methods and / or the key recognition methods based on key gaps provided in the embodiments of the present application.
[0143] Among them, the computer - readable storage medium may be an internal storage unit of the intelligent musical instrument described in the foregoing embodiment, such as the hard disk or memory of the intelligent musical instrument. The computer - readable storage medium may also be an external storage device of the intelligent musical instrument, such as a plug - in hard disk equipped on the intelligent musical instrument, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0144] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A piano key recognition method, characterized in that: The method comprises: S301, acquiring a first field of view image, wherein the first field of view image includes images corresponding to a plurality of piano keys of a type; S302, performing a first filtering operation on pixel points in the first field of view image based on pixel values in the first field of view image to suppress image noise of a first form, thereby obtaining a second field of view image; S303, performing preset threshold processing on the second field of view image to obtain a third field of view image, wherein the third field of view image includes a plurality of target pixel sets, and the target pixel sets are composed of a plurality of continuous pixel points whose pixel values are lower than a preset threshold; S304, performing a second filtering operation on the target pixel set in the third field of view image based on a preset structure element to suppress the second form of image noise, thereby obtaining a fifth field of view image; S305 , based on the second preset model, identifying a plurality of piano keys of the type from the fifth field of view image.
2. The method according to claim 1, characterized in that The S302 includes: Identify at least one pixel curve on the first field of view image, where the pixel curve is composed of a plurality of pixel points and is used to reflect a degree of change in pixel values of the pixel points; Calculating a pixel difference degree between at least one first target segment and an adjacent second target segment in the pixel curve, wherein the first target segment and the second target segment are composed of at least one pixel point; Determining whether the pixel difference degree exceeds a pixel difference threshold; If so, the first target segment is modified to reduce the pixel difference degree.
3. The method according to claim 1, characterized in that The second form of image noise includes burr noise and linear noise; S304 includes: Based on the first structure element, the boundary of each target pixel set in the third field of view image is extended to fill the internal holes or breaks of the target pixel set to obtain a fourth field of view image; Based on the second structure element, the boundaries of each target pixel set in the fourth field of view image are shrunk to eliminate the burr noise of each target pixel set and the linear noise in the fourth field of view image, so as to obtain the fifth field of view image.
4. The method according to claim 1, characterized in that: The method further comprises: Based on preset camera parameters, determining perspective deformation parameters of each of the keys of the type according to the position coordinates of the keys of the type in the fifth field of view image; The target shape parameters of each of the keys of the type are determined according to the reference shape parameters of the keys of the type and the perspective deformation parameters, and the contour of the corresponding keys of the type is adjusted according to the target shape parameters of each of the keys of the type.
5. The method according to claim 4, characterized in that The method further comprises: Acquire at least one proximal key of the first type whose perspective deformation parameter is lower than a first deformation threshold, and determine the reference morphological parameters of the first type of keys based on the morphological parameters of the at least one proximal key; Acquire at least one distal key of the type having the perspective deformation parameter higher than a second deformation threshold, wherein the first deformation threshold is less than or equal to the second deformation threshold; Adjusting the reference standard morphological parameters of the one type of keys based on the perspective deformation parameters of each of the one type of remote keys to determine the target morphological parameters of each of the one type of remote keys; According to the target morphological parameters of each of the distal keys, the contour of the corresponding distal keys is adjusted.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtain the actual shape parameters and position coordinate parameters of each type of piano key; Based on the preset key layout parameters, a plurality of the keys of the first category are screened according to the actual shape parameters and position coordinate parameters of the keys of the first category to determine a plurality of target keys of the first category; Output the position coordinates of a plurality of the target first-class piano keys in the first field of view image.
7. The method according to claim 1, characterized in that The preset key layout parameters include at least one of a target length, a target width, a target aspect ratio, a target area, a target key spacing, and a target key quantity of the key type.
8. An intelligent musical instrument, characterized in that: The intelligent musical instrument comprises: An image acquisition module, used for acquiring a first field of view image, wherein the first field of view image includes images corresponding to a plurality of first-class piano keys; Memory for storing computer programs; A processor, configured to execute the computer program and implement the key recognition method as claimed in any one of claims 1 to 7 when executing the computer program.
9. A computer device, characterized in that: The device comprises: Memory for storing computer programs; A processor, configured to execute the computer program and implement the key recognition method as claimed in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the key recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent playing error identification method and system for assisting piano teaching
CN113723264A
Key identification method and device, electronic equipment and storage medium
CN111695499A
Real-time visual key detection and positioning method for humanoid piano playing robot
CN114359314A
Intelligent identification method and system for giving assistance with piano teaching, and intelligent piano training method and system
WO2022052941A1