An electric piano touch dynamics calibration method and system based on AI adaptation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU RANTION TECH CO LTD
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]但现有电钢琴触键校准方法,多依赖于静态、通用的预设曲线或简单的手动调节,缺乏对演奏者个体能力差异、实时演奏语境及音乐情感意图进行系统性感知与动态融合的智能机制,不仅无法提供前瞻性、个性化的自适应校准引导,更可能因反馈滞后或建议失准而固化不良演奏习惯,制约演奏者艺术表现力的有效提升
[0004] This application provides an AI-adaptive electric piano key touch force calibration method and system to solve the above-mentioned technical problems.
Smart Images

Figure CN122529940A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of adaptive estimation, and in particular to an AI-adaptive method and system for calibrating the touch force of an electric piano. Background Technology
[0002] In the field of intelligent music education and personalized performance training, the electric piano, as a core interactive device, is a key indicator for measuring performance performance and skill proficiency through precise control of key pressure. It is directly related to the efficiency of skill advancement, the quality of musical emotional expression, and the immersive learning experience of the performer, and has become a bridge connecting traditional piano art and modern digital music technology.
[0003] However, existing electric piano touch calibration methods mostly rely on static, universal preset curves or simple manual adjustments. They lack intelligent mechanisms that systematically perceive and dynamically integrate individual differences in performers' abilities, real-time performance contexts, and musical emotional intentions. Not only do they fail to provide forward-looking, personalized adaptive calibration guidance, but they may also solidify bad playing habits due to delayed feedback or inaccurate suggestions, thus hindering the effective improvement of performers' artistic expression. Summary of the Invention
[0004] This application provides an AI-adaptive electric piano key touch force calibration method and system to solve the above-mentioned technical problems.
[0005] Firstly, this application provides an AI-adaptive method for calibrating the touch force of an electric piano, the method comprising: The system acquires a multi-source playing information set of the user, performs seamless collaborative identity recognition and initial state assessment based on the user's multi-source playing information set, and generates a user playing digital fingerprint information set; based on the user playing digital fingerprint information set, it performs multi-dimensional playing ability assessment and dynamic guidance strategy generation, and generates a personalized graded calibration strategy information set; based on the personalized graded calibration strategy information set, combined with real-time music context perception and emotional intent understanding, it performs forward-looking dynamic dynamic dynamic curve adjustment and optimization feedback generation, and generates and outputs a personalized touch force calibration suggestion report.
[0006] Through the above technical solutions, a fundamental shift from "one-size-fits-all" calibration to "personalized instruction" has been achieved. By constructing a user's digital fingerprint and understanding the emotional context of music, highly personalized, phased, and artistically relevant real-time guidance is provided, significantly improving learning efficiency and performance, and upgrading the digital piano into an intelligent music partner.
[0007] Optionally, the process of generating the user's performance digital fingerprint information set includes: acquiring multi-dimensional real-time playing data generated by the user's interaction with the electric piano through a mobile application client; determining the user's identity status based on the multi-dimensional real-time playing data using an intelligent recognition algorithm to construct a user identity identifier and its preliminary attribute tags; simultaneously, extracting initial key-touching behavior fingerprints based on the multi-dimensional real-time playing data to construct a set of playing features describing the user's basic playing habits and ability tendencies; and combining the user identity identifier and its preliminary attribute tags with the set of playing features to generate the user's performance digital fingerprint information set.
[0008] Optionally, the process of constructing the user identity identifier and its preliminary attribute tags includes: based on the historical storage data of the mobile application client, determining whether there is a verified account session initiated through the associated application; if so, the current session account information is directly used as the user identity identifier, and historical performance maturity tags are loaded from the cloud archive as the preliminary attribute tags; if not, an anonymous perception process is initiated, and the user's initial free playing segment is analyzed through AI feature comparison to extract its typical dynamic distribution and rhythmic stability features, and compared with the local temporary archive; if a match is found, the corresponding anonymous identifier and tag are used; if no match is found, a new anonymous identifier is created, and based on the typical dynamic distribution and rhythmic stability features, the preliminary attribute tags identifying the user's basic proficiency level are constructed.
[0009] Optionally, the process of constructing the playing feature set includes: extracting the basic dynamic distribution range of the user's continuous playing without guidance from the multi-dimensional real-time playing data, and analyzing the frequently used dynamic values and their dispersion; simultaneously extracting the user's typical rhythm control patterns, analyzing the time deviation rate when playing uniform notes, and the basic playing accuracy when handling simple rhythm patterns; and associating and encoding the basic dynamic distribution range, the frequently used dynamic values and their dispersion, the time deviation rate, and the basic playing accuracy to construct the playing feature set, which is used to characterize the user's original control habits and stability tendency without intervention.
[0010] Optionally, the process of generating the personalized graded calibration strategy information set includes: based on the playing feature set and the preliminary attribute tags, constructing a performance maturity profile containing specific intensity indicators and weak link identifiers through multi-dimensional performance ability assessment; based on the performance maturity profile, mapping and constructing a dynamic calibration rule set corresponding to the level through AI-driven three-level guidance strategy generation logic; and encapsulating the performance maturity profile and the dynamic calibration rule set to generate the personalized graded calibration strategy information set.
[0011] Optionally, the step of constructing a performance maturity profile through multi-dimensional performance ability assessment includes: based on the playing feature set, analyzing the user's control precision under different preset dynamic targets to generate a dynamic control stability score; analyzing the matching degree between the actual output dynamic range and the target range when the user processes musical phrases containing clear dynamic contrasts to generate a dynamic response awareness score; analyzing the clarity and evenness of notes when the user attempts basic techniques such as fast scales or arpeggios to generate a basic technique execution score; and combining the dynamic control stability score, the dynamic response awareness score, and the basic technique execution score with the preliminary attribute tags to construct the performance maturity profile represented by a radar chart.
[0012] Optionally, the AI-driven three-level guidance strategy generation logic constructs a dynamic calibration rule set, including: if the performance maturity profile indicates that the user is in the basic control ability formation stage, then an auxiliary shaping strategy rule is generated, the core rules of which include: identifying continuously too light key press intervals and generating visual feedback suggestions prompting the user to increase the pressing pressure, and identifying excessively heavy key presses that easily lead to tone breakage and generating auditory examples prompting the user to control the intensity; if the performance maturity profile indicates that the user is in the control ability expansion stage, then a collaborative reinforcement strategy rule is generated, the core rules of which include: when the user practices strong phrases, marking the difference between the peak intensity and the target peak intensity and providing dynamic range practice prompts, and when the user practices weak phrases, highlighting the intensity fluctuation intervals and providing stability control suggestions; if the performance maturity profile indicates that the user is in the stable and refined control stage, then a transparent resonance strategy rule is generated, the core rules of which include: recording the user's characteristic dynamic techniques in performance and comparing and analyzing them with standard interpretations to generate a descriptive feedback report for improving the refinement of artistic expression.
[0013] Optionally, the process of generating and outputting the personalized touch intensity calibration suggestion report includes: based on the real-time performance audio stream and score information, constructing an emotional context label describing the emotional tone of the current musical passage through real-time musical context perception; based on the emotional context label and the dynamic calibration rule set, generating instantaneous calibration parameters adapted to the current musical emotion through AI adaptive velocity curve tuning logic; based on the comparative analysis of the actual velocity in the entire performance passage and the expected velocity of the emotional context label, generating a text report containing specific improvement ranges and descriptive suggestions, and attaching the change history of the instantaneous calibration parameters as a report attachment; visually encapsulating and pushing the text report and the report attachment through the mobile application client to generate and output the personalized touch intensity calibration suggestion report.
[0014] Optionally, the process of constructing the emotional context label includes: using AI audio analysis technology to analyze the audio stream of the real-time performance, extracting the quantitative features of its chord tension and rhythmic regularity, and generating sound profile information that represents the musical atmosphere; simultaneously parsing the musical score information of the current performance segment, identifying the tempo terms, dynamic markings and expression terms, and converting them into clear emotional intensity and style description instructions; and through a pre-built AI music emotion mapping knowledge base, co-mapping and verifying the sound profile information with the emotional intensity and style description instructions to generate the emotional context label used to guide subsequent dynamic calibration.
[0015] Secondly, this application provides an AI-adaptive electric piano touch force calibration system, the system comprising: The performance digital fingerprint module is used to acquire a user's multi-source playing information set, and based on the user's multi-source playing information set, to perform identity-seamless collaborative identification and initial state assessment, generating a user performance digital fingerprint information set; the hierarchical calibration strategy module is used to perform multi-dimensional performance ability assessment and dynamic guidance strategy generation based on the user's performance digital fingerprint information set, generating a personalized hierarchical calibration strategy information set; the calibration suggestion report module is used to perform forward-looking dynamic dynamic dynamics curve adjustment and optimization feedback generation based on the personalized hierarchical calibration strategy information set, combined with real-time music context perception and emotional intent understanding, generating and outputting a personalized touch force calibration suggestion report. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram illustrating an application scenario provided in one embodiment of this application; Figure 2 A flowchart illustrating an AI-adaptive electric piano touch force calibration method provided in one embodiment of this application; Figure 3 This is a schematic diagram of an AI-adaptive electric piano touch force calibration system provided in one embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0019] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0020] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0021] Figure 1 This is a schematic diagram illustrating an application scenario provided by this application. In the process of calibrating the touch intensity of an electric piano, the method provided in this application can reduce bad playing habits ingrained due to delayed feedback or inaccurate suggestions, effectively improving the performer's artistic expression.
[0022] Specifically, the method of this application is applied to any server that communicates with the mobile application client, obtains the user's multi-source playing information set provided by the mobile application client through the server, and finally outputs a personalized key pressure calibration suggestion report to the pianist.
[0023] For specific implementation details, please refer to the following examples.
[0024] Figure 2 This is a flowchart illustrating an AI-adaptive electric piano key touch force calibration method according to an embodiment of this application. The method of this embodiment can be applied to the server in the above scenario. Figure 2 As shown, the method includes: S201. Obtain the user's multi-source playing information set, and based on the user's multi-source playing information set, perform identity-seamless collaborative identification and initial state assessment to generate the user's performance digital fingerprint information set.
[0025] The user multi-source playing information set is a data collection that comprehensively reflects a user's playing behavior and state when playing an electric piano. The data comes from the accompanying mobile application client (such as a dedicated app on a smartphone or tablet). It can collect ambient audio streams through the device's microphone, capture the user's hand posture through the camera, and access the user's historical practice records and preference settings.
[0026] Identity-insensitive collaborative identification and initial state assessment can be a processing logic that, without the user having any additional operational burden (i.e., "insensitive"), comprehensively utilizes the aforementioned multi-source data to intelligently identify the user's identity (regardless of whether they are logged into an account) and quantitatively assess their current basic performance ability.
[0027] The user's performance digital fingerprint information set can be the output of the above processing. It is a structured data set designed to create a unique and dynamically updated "performance identity profile" for each user. This profile not only contains identity identifiers, but more importantly, it accurately depicts the user's current original performance habits and ability baseline.
[0028] Specifically, digital piano touch force calibration is a crucial link connecting the player's physical movements with the digital sound source generation. Its core lies in establishing a mapping relationship between the user's finger pressure and the final output intensity and timbre. The goal is to finely adjust this mapping curve so that the digital piano can more realistically reproduce the touch response of a traditional piano and adapt to different players' control habits and skill levels, thereby improving performance expressiveness and control precision. However, existing digital piano calibration methods mostly rely on factory presets or manual global static adjustments by the user. A major limitation is that the calibration process lacks perception and modeling of the individual characteristics of the player, failing to distinguish the essential differences between beginners, intermediate players, and professional players. This results in a lack of targeted calibration suggestions and generic practice feedback. This solution addresses this problem by acquiring and analyzing a multi-source set of user playing information, performing seamless collaborative identification and initial state assessment, and generating a user playing digital fingerprint information set that uniquely identifies and quantifies the user's basic ability. This lays an irreplaceable data foundation for subsequent truly personalized adaptive calibration.
[0029] S202. Based on the user's digital fingerprint information set, perform multi-dimensional performance ability assessment and dynamic guidance strategy generation, and generate a personalized graded calibration strategy information set.
[0030] Multidimensional performance ability assessment can be a process of quantitatively scoring and analyzing a user's performance control ability from multiple professional dimensions based on the generated digital fingerprint (especially the set of playing features and preliminary attribute labels). These dimensions usually include, but are not limited to, dynamic control stability, dynamic response awareness, and basic skill execution accuracy, which together constitute a comprehensive ability evaluation system.
[0031] Dynamic guidance strategy generation can intelligently match and generate a set of strategies with clear guidance objectives and intervention rules based on the results of the above capability assessment.
[0032] The personalized graded calibration strategy information set can be the final output of the above evaluation and generation process. It is a structured strategy package that encapsulates finely graded calibration rules and guidance schemes for the current user.
[0033] Specifically, after obtaining accurate user digital fingerprints, generating effective guidance strategies is the core of achieving intelligent teaching. Traditional digital music teaching systems often only provide simple right-or-wrong judgments or general practice pieces. A major limitation is that the strategy generation process is disconnected from the user's real-time ability status, lacking a dynamic, graded guidance mechanism based on multi-dimensional ability assessment. Unlike human teachers, they cannot provide achievable challenges tailored to the student's current level, leading to low practice efficiency or increased frustration. This solution addresses this problem by using the user's performance digital fingerprint information set and employing an adaptive analysis engine to conduct multi-dimensional performance ability assessments. This constructs a performance maturity profile revealing the user's specific strengths and weaknesses, and dynamically generates guidance strategies accordingly. Ultimately, it produces a personalized, graded calibration strategy information set that precisely matches the user's current stage. This process ensures the scientific rigor and progressive nature of the guidance strategies, making it a crucial decision-making step for achieving individualized instruction.
[0034] S203. Based on the personalized graded calibration strategy information set, combined with real-time music context perception and emotional intent understanding, perform forward-looking dynamic force curve adjustment and optimization feedback generation, generate and output a personalized touch force calibration suggestion report.
[0035] Real-time music context perception and emotional intent understanding can refer to the process by which the system simultaneously analyzes the audio stream being played and the corresponding sheet music information during the user's performance, and judges in real time the emotional tone (such as cheerful, melancholic, or exciting), stylistic features (such as Baroque or Romantic), and artistic expression intent of the current musical passage.
[0036] Proactive dynamic velocity curve adjustment refers to the predictive fine-tuning of the electric piano's velocity-timbre response curve before or during the actual key press, based on an understanding of the current musical context and the user's personalized calibration strategy.
[0037] The personalized key pressure calibration suggestion report can be the final output of this method. It is not just a post-event data analysis, but a comprehensive feedback document that integrates real-time adjustment records, contextual performance analysis and specific improvement suggestions. It is usually presented to users in a visual form (such as charts, highlighted scores, and text descriptions) through mobile application clients.
[0038] Specifically, the highest level of performance calibration and guidance must serve the musical expression itself; dynamic control divorced from context is mechanical and meaningless. Another major limitation of existing technical solutions is that the calibration and feedback mechanisms are isolated from the musical work being performed, failing to understand the emotional tone, stylistic requirements, and artistic intent of musical phrases. This leads to suggestions that may violate the rules of musical expression or even solidify incorrect expressions. This solution addresses this deeper issue by combining the aforementioned personalized hierarchical calibration strategy information set with AI-driven real-time musical context perception and emotional intent understanding. During performance, it performs forward-looking dynamic dynamic curve adjustments, enabling calibration parameters to adapt to the emotional fluctuations of the music in real time. Ultimately, it generates a personalized touch dynamic calibration suggestion report that integrates contextual analysis, real-time adjustment records, and personalized improvement plans. This represents a fundamental leap from "mechanical parameter tuning" to "artistic expression assistance." The aforementioned context perception and report generation process can be implemented as an independent music intelligence service and report generation service within the system's microservice architecture, working collaboratively to achieve real-time analysis and comprehensive feedback output.
[0039] The method provided in this embodiment first uses the digital piano's sensors and a mobile application to seamlessly collect multi-source playing data from the user. Through identity collaborative identification and feature extraction, a dynamically updated digital fingerprint of the user's performance is constructed. Then, based on this fingerprint, a multi-dimensional ability assessment is performed, generating a graded calibration strategy that precisely matches the user's current level. Finally, during real-time performance, the system synchronously analyzes the musical emotional context and makes forward-looking dynamic fine-tuning of the dynamic response curve accordingly. Simultaneously, a calibration report integrating contextual understanding and personalized suggestions is generated, completing a closed loop from perception and diagnosis to adaptive intervention. This solution achieves a fundamental shift from "one-size-fits-all" calibration to "personalized instruction." By constructing a user's digital fingerprint and understanding the musical emotional context, it provides highly personalized, phased, and artistically aligned real-time guidance, significantly improving learning efficiency and performance expressiveness, upgrading the digital piano into an intelligent music partner.
[0040] In some embodiments, multidimensional real-time playing data generated by the user's interaction with the electric piano is acquired through a mobile application client. Based on the multidimensional real-time playing data, an intelligent recognition algorithm is used to determine the user's identity status and construct a user identity identifier and its preliminary attribute tags. At the same time, based on the multidimensional real-time playing data, the initial key-touching behavior fingerprint is extracted to construct a set of playing features describing the user's basic playing habits and ability tendencies. The user identity identifier and its preliminary attribute tags are combined with the set of playing features to generate a set of digital fingerprint information of the user's performance.
[0041] In this step, the intelligent recognition algorithm specifically refers to a pre-defined feature comparison model based on deep learning. Its input is a fragment of the user's real-time playing data, and its output is a similarity score between the fragment and a known user profile, or a feature vector used to characterize the uniqueness of an anonymous user. The core of this algorithm is to extract stable and distinctive personal habit patterns from limited, unconstrained free playing.
[0042] A user identifier and its initial attribute tags can be a structured data pair; the user identifier is a string (such as an account ID or anonymous UUID) used internally by the system to uniquely track users; the initial attribute tags are metadata bound to the identifier that describes the user's basic proficiency level.
[0043] The playing feature set can be a structured set of values output from the initial key-touching behavior fingerprint extraction process. It encapsulates specific dimensions of measurement values such as mean velocity, standard deviation of velocity, and rhythm deviation rate. It is a core component of the user's playing digital fingerprint information that centrally describes the behavioral pattern.
[0044] Specifically, traditional digital music systems suffer from fragmented user identification: either forcing login and interrupting the experience, or treating each practice session as an independent event in an anonymous state, failing to create a coherent personal growth profile. This step, through the parallel processing of seamless collaborative identity recognition and initial touch behavior fingerprint extraction, fundamentally solves the challenge of "seamlessly establishing a continuous personalized service baseline." For example, when an anonymous user uses the same keyboard for the second time, the system can automatically recognize and continue their historical progress without any manual intervention, a prerequisite for truly adaptive learning.
[0045] In the specific analysis process, the mobile application client continuously receives Note-On / Note-Off events (including velocity values, pitch, and timestamps) sent by the electric piano via the Bluetooth MIDI protocol, and simultaneously records the ambient audio stream. For identity verification: the system first checks if an active OAuth session token exists locally. If it exists, it directly retrieves the user_id and historical probability_tag of the account from the cloud to complete the verification. If it does not exist, the anonymization process is initiated: the client encapsulates the user's free-playing MIDI data from the previous 30 seconds (e.g., a C major scale or arbitrary chords) into a data packet and calls the locally deployed intelligent recognition algorithm (a lightweight pre-trained Siamese neural network). This network extracts the feature vector of this data and performs a cosine similarity comparison with the anonymous user feature vector stored in the local SQLite database. If the highest similarity exceeds the threshold of 0.7, it is determined to be a returning user, and the corresponding anonymous UUID and tag are returned; otherwise, a new UUID is generated, and the initial touch behavior fingerprint extraction module is called. This module analyzes the same MIDI data: it calculates a histogram of velocity values for all Note-On events to obtain the most frequent velocity value (mode_velocity) and its standard deviation (std_velocity); it also calculates the coefficient of variation of the time interval between consecutive notes as timing_instability. These feature values constitute the playing feature set. Finally, the system packages the determined (user_id or anonymous_uuid plus probability_tag) with the calculated playing feature set to generate a JSON-formatted digital fingerprint information set of the user's performance.
[0046] In alternative or modified implementations, the mobile application client can be replaced by an embedded system or desktop software. The intelligent recognition algorithm can be replaced by a traditional machine learning model based on random forests or support vector machines (SVMs), whose preset process relies on offline training using a large amount of previously collected labeled user playing data. The rules for initial keystroke behavior fingerprint extraction can be extended, for example, by adding analysis of pedal usage frequency, or by using more complex temporal models (such as hidden Markov models) to capture patterns of dynamic changes. The local anonymous archive can also be replaced by a simple association method based on device fingerprints (such as Bluetooth MAC address hashes) as a fallback solution.
[0047] In some embodiments, based on historical stored data from the mobile application client, it is determined whether a verified account session initiated through the associated application exists: if it exists, the current session account information is directly used as the user's identity identifier, and historical performance maturity tags are loaded from the cloud archive as preliminary attribute tags; if it does not exist, an anonymous perception process is initiated, and the user's initial free playing segment is analyzed through AI feature comparison to extract its typical dynamic distribution and rhythmic stability features, and compared with the local temporary archive: if they match, the corresponding anonymous identifier and tag are used; if they do not match, a new anonymous identifier is created, and preliminary attribute tags identifying its basic proficiency level are constructed based on the typical dynamic distribution and rhythmic stability features.
[0048] Anonymous awareness process can be a set of non-intrusive recognition logic based on behavioral biometrics (playing habits) that the system initiates to identify returning users when there is no verified account session. Its core is AI feature comparison, which aims to associate the current anonymous user with all anonymous users in history to achieve continuity across usage.
[0049] A local temporary archive can be a lightweight database stored locally on the mobile device (such as an SQLite database) to cache the playing feature sets of all anonymous users and their corresponding anonymous identifiers and preliminary attribute tags. Its lifecycle is usually consistent with the application installation cycle.
[0050] Specifically, traditional solutions present a binary choice in user identity processing: forced login disrupts the user experience, while complete anonymity loses user continuity. This step, through a dual-track system of "session verification first, with anonymous AI comparison as a fallback," fundamentally resolves the contradiction between "continuity" and "seamlessness" in identity recognition. For example, even if a user has never logged in, when they practice on the same device a second time, the system can still "recognize" them and retrieve their previous progress, which is the foundation for building trustworthy personalized services.
[0051] In the specific analysis process, when the system needs to construct a user identity and its preliminary attribute tags, it first performs a session check. The mobile application calls the secure storage API provided by the operating system (such as Keychain for iOS or Keystore for Android) to query whether a valid OAuth access token exists. If it exists, the token is used to send an HTTPS GET request to the cloud server to the / api / user / profile endpoint. After verifying the token's validity, the server returns a JSON response containing user_id and historical_performance_tag. The system directly uses this user_id as the user identity and the tag returned by the cloud as the preliminary attribute tag. If no valid session exists, the anonymity awareness process is initiated: the system extracts the current user's initial free-playing segment (such as the first 60 seconds of MIDI data) and generates a 128-dimensional feature vector using a feature extractor (such as converting the velocity sequence into a standardized histogram). Then, the AI feature comparison module is called: this module is pre-configured as a local vector index built based on the Faiss (Facebook AI Similarity Search) library. The system inputs the new feature vector into the index, performs a nearest neighbor search (K=1), and returns the ID (i.e., anonymous UUID) of the most similar vector and its similarity score. If a match is found (similarity > 0.75), the ID is used as the user's identifier, and its associated tag is read from the local temporary archive as the initial attribute label. If no match is found, the system generates a new UUID (e.g., using java.util.UUID.randomUUID()), and based on the features of the current segment, determines its basic proficiency level through a preset decision tree rule set (e.g., if the average intensity is < 40 and the rhythm deviation is > 20%, then the level is "beginner"), forming the initial preliminary attribute label, and stores the new ID, feature vector, and label together in the local archive.
[0052] In alternative or modified implementations, the source of verified account sessions can be extended to device-bound login via scanning a QR code on the digital piano body. The cloud archive's data structure can include more granular capability history details. The default AI feature matching algorithm can be replaced with an approximate matching algorithm based on Locality Sensitive Hash (LSH) for faster retrieval on mobile devices; the default hash function and bucket size need to be determined during development. The local temporary archive can employ an LRU (Least Recently Used) strategy for capacity management, automatically cleaning up the least recently used anonymous archives.
[0053] In some embodiments, the basic dynamic distribution range of a user playing continuously without guidance is extracted from multidimensional real-time playing data, and the frequently used dynamic values and their dispersion are analyzed. The user's typical rhythm control pattern is extracted simultaneously, and the time deviation rate when playing uniform notes and the basic playing accuracy when handling simple rhythm patterns are analyzed. The basic dynamic distribution range, frequently used dynamic values and their dispersion, time deviation rate and basic playing accuracy are correlated and encoded to construct a playing feature set. The playing feature set is used to characterize the user's original control habits and stability tendency without intervention.
[0054] The basic velocity distribution range can be a numerical range [P10, P90] obtained by statistically analyzing the velocity values (Velocity, usually 1-127) of all notes in a continuous, unguided MIDI playing data from a user, and by calculating preset percentiles (such as the 10th percentile P10 and the 90th percentile P90).
[0055] Frequent use of velocity values can be determined by counting the velocity values (typically ranging from 1 to 127) of all notes in a continuous MIDI recording played by the user, and then calculating the mode, which is the velocity value that appears most frequently.
[0056] Dispersion can refer to the standard deviation of the same set of force values. This value quantifies the range of fluctuation of a user's touch force around their habitual value (frequently used force value). The larger the standard deviation, the more unstable and fluctuating the user's force control is; the smaller the standard deviation, the more concentrated and stable the force control is.
[0057] The time deviation rate can be the average relative error of the actual inter-onset interval (IOI) relative to the target time value when playing a series of notes with theoretically uniform time values (such as a scale on a metronome).
[0058] Basic playing accuracy can be measured by the system's Dynamic Time Warping (DTW) algorithm when a user attempts to play a preset simple rhythmic pattern (such as "quarter note - two eighth notes"). The system aligns and matches the actual IOI sequence played with the IOI template of the target rhythmic pattern and calculates the normalized cumulative distance. The smaller this value, the higher the accuracy of rhythmic imitation.
[0059] Specifically, traditional performance analysis often only records raw data from a single performance or makes simple right-or-wrong judgments, failing to extract the user's stable, intrinsic control habits. This step, by defining and extracting quantitative features such as "dynamic distribution range" and "time deviation rate," fundamentally solves the problem of how to "objectively characterize the user's original muscle memory and neural control patterns." For example, it can clearly distinguish between a user who is "habitually gentle but rhythmically unstable" and a user who is "bold in dynamics but rhythmically precise," providing irreplaceable objective evidence for subsequent "personalized instruction."
[0060] In the specific analysis process, the system acquired approximately 60 seconds of multidimensional real-time playing data (MIDI event stream) from the user's unguided playing. First, the velocity value list V_list of all Note-On events was extracted. The 10th and 90th percentiles of V_list were calculated using NumPy's percentile function to obtain the basic velocity distribution interval [v_p10, v_p90]. The mode of V_list was calculated using SciPy's mode function as the frequently used velocity value v_mode, and the standard deviation was calculated using the std function as the dispersion v_std. Next, the press timestamps of consecutive notes in the same data segment were extracted, and the time difference between adjacent notes was calculated to obtain the actual IOI sequence IOI_actual. Let the target time value be T (e.g., 500ms following a metronome), and the time deviation rate TDR = np.std(IOI_actual) / T * 100. Then, the system prompted the user to imitate a preset simple rhythmic pattern (e.g., "dong-da-da"), obtaining its IOI sequence IOI_user. The pre-defined DTW algorithm (using the fastdtw library) is invoked to calculate the shortest path distance dtw_dist between IOI_user and the standard template IOI_template. This distance is then normalized to the [0,1] interval to obtain the basic playing accuracy score acc_score (the smaller the distance, the higher the score). Finally, through associative encoding, the five values (v_p10, v_p90, v_mode, v_std, TDR, acc_score) are written into a Python dictionary as key-value pairs, thus completing the construction of the playing feature set.
[0061] In alternative or modified implementations, the statistical method for determining the basic dynamic range can be replaced by using kernel density estimation (KDE) to identify the region with the highest probability density. Identification of frequently used dynamic values can be replaced by using cluster centers derived from K-Means clustering. The calculation of the time deviation rate can incorporate more robust statistics, such as median absolute deviation (MAD). The DTW algorithm for assessing basic playing accuracy can be replaced by sequence matching based on a hidden Markov model (HMM), which presupposes training the model with a large amount of standard rhythm performance data. The library of simple rhythm patterns can be expanded, and different difficulty levels of rhythm patterns can be dynamically selected for testing based on the user's skill level.
[0062] In some embodiments, based on the set of playing features and preliminary attribute labels, a performance maturity profile containing specific intensity indicators and weak point identifiers is constructed through multi-dimensional performance ability assessment; based on the performance maturity profile, a dynamic calibration rule set corresponding to the level is mapped and constructed through an AI-driven three-level guidance strategy generation logic; the performance maturity profile and the dynamic calibration rule set are encapsulated to generate a personalized graded calibration strategy information set.
[0063] Multidimensional performance ability assessment can be a pre-set comprehensive assessment module. Its input is a set of playing features and preliminary attribute labels, and its output is a set of specific scores covering different performance dimensions and weak points identified based on these scores (such as "weakness: weak playing stability").
[0064] A performance maturity profile is a structured data object output by a multidimensional performance ability assessment module. It includes not only specific intensity indicators (scores) for each dimension, but also a qualitative assessment of the user's overall stage (such as "the stage of forming basic control ability") and a list of weak points.
[0065] The three-level guidance strategy generation logic can be a pre-defined decision logic mapper. Its core is a rule-based expert system or a simple classification model. It takes the qualitative judgment about the user's stage in the performance maturity profile as input, maps it to one of the three predefined strategy levels, and triggers the rule generator of the corresponding level to output a set of specific, executable guidance rules.
[0066] The dynamic calibration rule set can be a structured list of "condition-action" rules generated by the above strategy.
[0067] Specifically, traditional music teaching software or functions typically offer fixed, uniform practice courses, unable to dynamically adjust to users' ever-changing skill gaps. This step addresses the core challenge of "how to transform diagnosis into immediate, actionable, personalized teaching actions" by mapping objective assessment results (profiles) to differentiated teaching strategies (rule sets) in real time. For example, the system can automatically generate practice suggestions focusing on dynamic gradations for a beginner with a strong sense of rhythm but weak dynamic control, rather than having them repeatedly practice generic scales.
[0068] In the specific analysis process, the system first performs a multi-dimensional performance ability assessment: it calls a pre-defined evaluation function, `evaluate_maturity(profile)`, which reads fields such as `velocity_std` (stress standard deviation) and `timing_deviation_rate` (timing deviation rate) from the playing feature set. It calculates percentage scores for three dimensions—"dynamic control stability," "dynamic response awareness," and "basic skill execution"—using an internally pre-defined linear mapping formula (e.g., stability score = 100 - velocity_std * 2) and threshold comparisons (e.g., if `timing_deviation_rate > 20%`, then rhythmic stability is marked as a weak point). It then generates a list of weak point identifiers. These results collectively constitute a performance maturity profile, typically formatted as a dictionary containing {"scores": [...], "weaknesses": [...]}, which can be rendered as a radar chart by the visualization module. Subsequently, the system inputs the performance maturity profile into an AI-driven three-level guidance strategy generation logic. This logic is a pre-defined rule lookup table implemented using a Python dictionary. For example, its core rule is: if all(scores<60): stage = “basic”; elif all(scores>80): stage = “advanced”; else:stage = “intermediate”. Based on the determined stage value, the system loads the corresponding dynamic calibration rule set (a list of if-condition: then-action rules) from three preset independent rule bases (corresponding to the “assisted shaping”, “synergistic enhancement”, and “clear resonance” stages respectively). Finally, the system packages the performance maturity profile and the dynamic calibration rule set into a JSON object, i.e., a personalized graded calibration strategy information set, and stores it in memory or local cache for subsequent real-time calibration module to call.
[0069] In alternative or modified implementations, the multidimensional performance ability assessment model can be replaced with a classifier based on a lightweight neural network (such as a multilayer perceptron, MLP), which requires supervised training through the collection of a large amount of labeled user playing data. The AI-driven three-level guidance strategy generation logic can be upgraded to a more complex inference engine based on decision trees or Bayesian networks, whose preset rules can be automatically generated by analyzing teaching records from excellent teachers. The dynamic calibration rule set can be described in YAML or XML format for easy editing and updating by non-technical personnel. The grading strategy can be expanded from three to five levels to provide more granular guidance.
[0070] In some embodiments, based on the set of playing features, the system analyzes the user's control precision under different preset dynamic targets to generate a dynamic control stability score; it analyzes the matching degree between the actual output dynamic range and the target range when the user processes musical phrases containing clear dynamic contrasts to generate a dynamic response awareness score; it analyzes the clarity and evenness of notes when the user attempts basic techniques such as fast scales or arpeggios to generate a basic technique execution score; and it combines the dynamic control stability score, dynamic response awareness score, and basic technique execution score with preliminary attribute labels to construct a performance maturity profile represented by a radar chart.
[0071] Preset dynamic targets refer to specific dynamic values that the system requires users to attempt during assessment exercises. These targets are usually set up as a list based on the dynamic markings of common etudes (such as p approximately 40, mf approximately 70, f approximately 90) and teaching experience. For example, targets = [30, 50, 70] means that users are required to play at three different dynamic levels: soft, medium, and loud.
[0072] Clarity can refer to the precision of the attack (start) of each note and the degree of separation between notes when playing rapidly and continuously. In technical assessment, it can be quantified by analyzing the energy rise slope of the starting segment of each note or the overlap of the time values of adjacent notes.
[0073] Uniformity can refer to the consistency of the dynamics and duration of each note when playing a series of notes in succession. It can be quantified by calculating the standard deviation of the dynamic values of these notes and the coefficient of variation (CV) of the tempo interval (IOI).
[0074] Specifically, traditional assessments often provide vague evaluations like "good" or "needs improvement," leaving users unable to pinpoint the specific issue: inaccurate dynamic control, lack of contrast between loud and soft passages, or unclear technique. This step addresses the core problem of "vague and ineffective training guidance" in assessment results by breaking them down into three quantifiable specific scores. For example, it can clearly indicate that the user's "dynamic response awareness (65 points) is significantly lagging behind technique execution (80 points)," thus precisely guiding the practice focus from blindly increasing speed to paying attention to dynamic changes in musical phrases.
[0075] In the specific analysis process, the system first retrieves the corresponding assessment task from a pre-set practice library based on the level indicated in the initial attribute labels. For the dynamic control stability score, the system instructs the user to play the same set of scales using three different pre-set dynamic targets (e.g., 30, 60, 90). The system records the actual dynamic value sequence for each performance, calculates the standard deviation of each sequence, and then calculates the average of these three standard deviations. Finally, it converts the result into a percentage score using a pre-set linear mapping function (e.g., stability score = 100 - average standard deviation * 1.5). For the dynamic response awareness score, the system provides a test phrase clearly marked with p (weak) and f (loud) contrasts. After the user plays, the system extracts the actual average dynamic range of the p and f sections of the phrase and calculates the actual output dynamic range achieved by the user (average force of f section - average force of p section). The actual range is compared with the target range (preset, e.g., 90-40=50), and the matching percentage is calculated as: Matching Percentage = (Actual Range / Target Range) * 100. This matching percentage is then used directly as the score (anything over 100% is counted as 100%). For basic skill performance scoring, the user is asked to quickly play a C major scale. The system analyzes the recording or MIDI data: Clarity is assessed by calculating the average rise slope of energy within 10ms after the start of each note; Evenness is assessed by calculating the standard deviation of all note velocity values and the coefficient of variation of rhythmic IOI. These two indicators are scored separately and then weighted to obtain the final skill performance score. Finally, the system inputs the three scores and preliminary attribute labels (e.g., "Adult Beginner") into a report generation template, calls radar chart plotting functions such as those in the matplotlib library, and automatically generates a radar chart-based portrait image file representing the performance maturity.
[0076] In alternative or modified implementations, the assessment of dynamic control stability can be replaced by requiring the user to play along a dynamically changing dynamic target curve (such as a crescendo-diminuendo curve), using the root mean square error (RMSE) as the scoring criterion. The target range for dynamic response awareness scoring can be dynamically adjusted based on style databases from different historical periods (Baroque, Romantic). The basic skill task of assessing clarity and evenness can be replaced by scales with arpeggios or vibrato. The radar chart representation can be extended to a star chart that includes more dimensions (such as rhythmic stability and pedal use), and the scoring mapping function can be non-linearly normalized using a piecewise function or a sigmoid function based on a large number of sample statistics.
[0077] In some embodiments, if the performance maturity profile indicates that the user is in the stage of forming basic control ability, then auxiliary shaping strategy rules are generated. The core rules include: identifying continuously too light key press intervals and generating visual feedback suggestions prompting the user to increase the pressing pressure, and identifying excessively heavy key presses that easily lead to tone breakage and generating auditory examples prompting the user to control the intensity. If the performance maturity profile indicates that the user is in the stage of expanding control ability, then collaborative reinforcement strategy rules are generated. The core rules include: when the user practices forte phrases, marking the difference between the peak intensity and the target peak intensity and providing dynamic range practice prompts; when the user practices forte phrases, highlighting the intensity fluctuation intervals and providing stability control suggestions. If the performance maturity profile indicates that the user is in the stage of stable and refined control, then transparent resonance strategy rules are generated. The core rules include: recording the user's characteristic dynamic techniques in performance and comparing and analyzing them with standard interpretations to generate descriptive feedback reports for improving the refinement of artistic expression.
[0078] The auxiliary shaping strategy rules can be a pre-set rule base for users who are in the basic control ability formation stage (usually referring to the overall score being less than 60 points). Its core goal is to establish the correct concept of dynamic range and prevent the formation of extremely wrong key touch habits. Each rule in the rule base contains a trigger condition (such as "N consecutive notes with dynamics < threshold L") and a feedback action (such as "highlighting and displaying prompt text T at the corresponding position in the score").
[0079] Collaborative reinforcement strategy rules can be a pre-set rule base for users in the control ability expansion stage (usually referring to a comprehensive score between 60 and 80 points). Its core goal is to help users consciously use and consolidate dynamic contrast in musical performance. The rules focus on quantitative analysis and prompts for the performance of "strong" (f) and "weak" (p) in specific musical phrase contexts.
[0080] The Transparent Resonance Strategy Rules can be a pre-set rule library for users in a stable and refined control stage (usually referring to a comprehensive score of over 80 points). Its core goal is to enhance the subtlety and artistic individuality of musical expression. The rules no longer focus on basic right or wrong, but rather analyze the compatibility between the user's unique dynamic processing techniques (such as personalized crescendo and diminuendo curves) and the musical emotions.
[0081] Specifically, traditional automated teaching systems typically employ a "one-size-fits-all" approach with fixed exercises or generic feedback, failing to provide dynamic, evolving guidance strategies based on the user's skill progression from "beginner" to "intermediate user" and then to "performer." This step addresses the core issue of "how to ensure the teaching system's intelligence evolves in sync with the user's growth" by pre-setting three sets of AI-driven, three-tiered guidance strategies with distinct goals, methods, and feedback formats. For example, it ensures that beginners focus only on "not being too light or too heavy," intermediate users are guided on "how to make clear contrasts between strengths and weaknesses," and advanced users are explored on "how to make progressive overcoming more impactful."
[0082] In the specific analysis process, the system first reads the three-dimensional scores from the performance maturity profile. It executes a preset judgment logic: if all three scores are below the threshold θ1 (e.g., 60), it is judged as "basic stage"; if all are above the threshold θ2 (e.g., 80), it is judged as "refined stage"; otherwise, it is "expansion stage". After the judgment, the system loads the corresponding preset rule library (e.g., basic_rules.json) from local or cloud. For the basic stage, the system applies auxiliary shaping strategy rules: it monitors the note flow in real time, and when it detects that the dynamics of three consecutive notes are less than the "too light threshold" (e.g., 20), it immediately triggers the rule, overlaying a flashing "↑" icon and a visual feedback suggestion of "please increase the dynamics" on the screen at the positions of these notes; at the same time, when it detects that the dynamics of a single note exceed the "too heavy threshold" (e.g., 110), another rule is triggered, and in addition to the visual cue, the audio engine is immediately called to play an auditory example of the same note with a dynamics of 90. For the expansion phase, the system applies a collaborative reinforcement strategy: when a user practices a phrase marked with 'f', the system records the peak dynamics of that phrase and compares it with the target peak (e.g., 90). If the difference is greater than 10, a dynamic range practice prompt pops up in the sidebar after the phrase ends, such as "The peak dynamics of the forte phrase is XX, the target is 90, please try using more arm strength." For the refinement phase, the system applies a transparent resonance strategy: after a user plays a complete piece, the system performs dynamic time warping (DTW) alignment and difference analysis on its entire dynamics curve against a pre-stored "standard interpretation" dynamics curve, generating a descriptive feedback report, pointing out, for example, "You started the crescendo a little too early in the fifth measure, causing the peak to sound slightly rushed."
[0083] In alternative or modified implementations, the logic for stage determination can be upgraded from simple threshold comparison to a classifier based on Naive Bayes or a small neural network, whose preset model requires training with a large amount of labeled user profile data. The rule base can be constructed semi-automatically using natural language processing (NLP) and pattern mining techniques by analyzing lesson plans and real-time instruction records from experienced piano teachers. Visual feedback suggestions can be enhanced in AR (augmented reality) form, projecting light spots above real piano keys. Auditory examples can be provided not singly, but in multiple versions by different performers for users to listen to and compare. The generation of descriptive feedback reports can integrate a large language model (LLM) to make the language more natural and inspiring.
[0084] In some embodiments, based on the real-time audio stream and score information of the performance, an emotional context label describing the emotional tone of the current musical passage is constructed through real-time musical context perception. Based on the emotional context label and a set of dynamic calibration rules, instantaneous calibration parameters adapted to the current musical emotion are generated through AI adaptive dynamic curve tuning logic. Based on a comparative analysis of the actual dynamics and the expected dynamics of the emotional context label throughout the entire performance passage, a text report containing specific improvement ranges and descriptive suggestions is generated, and the change history of the instantaneous calibration parameters is attached to the report. The text report and the report attachment are visualized and pushed through a mobile application client to generate and output a personalized touch dynamics calibration suggestion report.
[0085] Real-time music context awareness can refer to a pre-set AI analysis module that processes audio streams (the sound of the user's actual performance) and sheet music information (standard MIDI or MusicXML files) in parallel, infers the musical emotional tone of the current performance segment in real time (such as "cheerful", "sad", "exhilarating"), and outputs structured emotional context labels.
[0086] Emotional context labeling can be a structured data object, which is the output of real-time music context perception. It contains at least two fields: "emotional dimension" (such as joy, sadness, tension) and "intensity level" (such as low, medium, high), which are used to quantitatively guide the direction of subsequent dynamic calibration.
[0087] AI adaptive dynamic curve tuning logic can be a preset dynamic parameter adjustment engine. It takes emotional context labels and the current user's dynamic calibration rule set as input. Its core is a set of "context-parameter" mapping functions. This logic can calculate specific and safe adjustment parameters in real time based on the musical emotion (such as "exhilarating" which requires an overall increase in dynamics) and the user's current ability stage (such as "basic stage" which requires limiting the adjustment range).
[0088] Instantaneous calibration parameters can be a set of values output by the tuning logic described above, which can be immediately applied to the electric piano sound source or haptic engine. They are usually a data structure containing fields such as "basic velocity offset" (e.g., +5), "dynamic range scaling factor" (e.g., 1.2), and "velocity response curve steepness".
[0089] Specifically, traditional electric piano dynamics calibration suggestions often focus on correcting technical errors (such as "playing too softly here"), completely ignoring the core of musical expression—emotional expression—making practice mechanical and tedious. This step, by introducing real-time musical context awareness and AI-adaptive dynamics curve tuning logic, fundamentally solves the problem of "how to combine cold, technical calibration with vibrant musical emotional expression." For example, when a user plays a melancholic nocturne, the system not only corrects wrong notes but also suggests "you can control the dynamics here more gently to enhance the tranquil atmosphere," achieving a leap from "correct playing" to "expressive performance."
[0090] In the specific analysis process, when a user begins playing a piece of music, the system launches two parallel threads: Thread one calls the real-time music context awareness module to extract chord and rhythm features from the incoming audio stream, while simultaneously parsing the expression markings (such as dolce, agitato) in the score information of the current measure; the module generates an emotional context label (such as {"emotion": "agitated", "intensity": "high"}) every 2-4 measures using a pre-set classification model trained on music psychology data. Thread two loads a set of dynamic calibration rules applicable to the user (such as collaborative reinforcement strategy rules). Subsequently, the AI adaptive velocity curve tuning logic is activated: it first reads the emotional context label and queries the pre-set "emotion-base offset" mapping table (such as "agitated" corresponding to a base offset + 8); then, based on the user stage indicated by the set of dynamic calibration rules, it performs a "safe" scaling of the offset (for example, a scaling factor of 0.5 for the base stage user and 1.2 for the fine stage user), ultimately generating a set of instantaneous calibration parameters and fine-tuning the dynamic response curve of the electric piano in real time. After the performance, the system performs offline analysis: it compares the dynamic sequence of each note actually played by the user with the expected dynamic curve (a "target line" fluctuating according to the expression markings in the score) dynamically generated based on emotional context tags, and calculates the root mean square error (RMSE). Based on the comparison results and feedback templates in the dynamic calibration rule set, the analysis engine automatically generates a text report, such as "Your dynamic peak reached 85 in the exciting passage (measures 5-8), which is very good; however, the dynamic fluctuation was large in the gentle connecting phrase (measure 12) (standard deviation 12), and it is recommended to practice slow touch to improve control." At the same time, the system generates a time-series diagram of the instantaneous calibration parameter changes during the performance, as an attachment to the report. Finally, the report rendering module of the mobile application client packages the text and charts into an interactive H5 page, which is pushed to the user through the in-app messaging system.
[0091] In alternative or modified implementations, real-time music context awareness can be based on more sophisticated music information retrieval (MIR) techniques, such as using convolutional neural networks (CNNs) to directly analyze the Mel spectrogram of audio to identify emotions. The mapping rules in the AI-adaptive dynamics curve tuning logic can be non-static lookup tables, but rather a small reinforcement learning (RL) agent that continuously optimizes the adjustment strategy through user interaction. The desired dynamics curve can be non-single-standard, but rather provide a "desired dynamics band" composed of versions from multiple performers. Report generation can integrate a large language model (LLM) to make the text reports more personalized and encouraging. Push notifications can be delivered via email or synchronized with a teacher collaboration platform.
[0092] In some embodiments, AI audio analysis technology is used to analyze the audio stream of a real-time performance, extract the quantitative features of its chord tension and rhythmic regularity, and generate sound profile information that represents the musical atmosphere; the score information of the current performance section is parsed simultaneously, identifying tempo terms, dynamic markings and expression terms, and converting them into clear emotional intensity and style description instructions; through a pre-built AI music emotion mapping knowledge base, the sound profile information and emotional intensity and style description instructions are collaboratively mapped and verified to generate emotional context labels for guiding subsequent dynamic calibration.
[0093] Chord tension is a quantitative music theory index that is calculated from real-time performance audio streams using AI audio analysis technology (such as pre-trained neural network models). It is used to objectively describe the auditory tension caused by chords or harmonic progressions. The higher the value, the stronger the tension.
[0094] Rhythmic regularity can be calculated by comparing the deviation between the start time of the actual played notes and the standard beat point in the score. It is a statistical quantitative indicator (such as root mean square error or coefficient of variation) used to objectively measure the stability and accuracy of the performance rhythm. The lower the value, the more regular the rhythm.
[0095] Sound profile information can be a multi-dimensional feature vector composed of chord tension, rhythmic regularity, and other audio features (such as average loudness and spectral centroid). It is a digital and structured summary of the overall auditory characteristics of the current performance segment.
[0096] Tempo terms can be markers (such as Adagio, Allegro) parsed from musical score information that indicate the speed of performance, and are converted in the system into quantifiable values (such as beats per minute, BPM) or category labels that can be used for calculation.
[0097] Dynamic markings can be symbols (such as pp, mf, f) that are parsed from musical score information and indicate the loudness or softness of notes or passages. In the system, they are converted into corresponding reference dynamic target values or dynamic ranges.
[0098] Expression terms can be textual markers (such as dolce, espressivo) that are parsed from musical score information and indicate musical emotions or performance styles. These are converted into structured style description tags in the system.
[0099] The emotional intensity and style description instruction can be a structured data object output by the music score parsing module. It integrates the semantics of tempo terms, dynamic markings, and expression terms, and includes fields such as "emotional intensity" (e.g., high, low) and "style description" (e.g., lively, mellow) to guide emotional mapping.
[0100] A pre-built AI music emotion mapping knowledge base can be a model or knowledge system pre-built using machine learning methods (such as training on a large amount of labeled music data). Its core function is to fuse, map, and verify the sound profile information with the emotional instructions of the musical score, and finally output a unified and reliable emotional context label.
[0101] Specifically, traditional sentiment analysis often relies solely on musical notation or audio features, severing the connection between the "composer's intention" (musical score) and the "performer's realization" (sound), which easily leads to misjudgments. For example, the score may be marked as "sad," but if the performer's rhythm is unstable and the sound is noisy due to technical limitations, the score alone may incorrectly maintain the "sad" label, and the audio alone may misjudge it as "panic." This step, through collaborative mapping and verification, fundamentally solves the problem of "how to synthesize the guiding intent of the score with the actual sound effect of the performance to obtain the most accurate emotional tone of the current musical moment," providing a reliable contextual basis for subsequent precise dynamic calibration.
[0102] In the specific analysis process, the system initiates two parallel feature extraction pipelines. Pipeline A (Audio Analysis): The real-time audio stream is processed by frame segmentation (e.g., one frame every 100ms). Open-source audio processing libraries (e.g., librosa) are used to extract chord features for each frame, which are then input into a pre-trained chord tension calculation model (e.g., an LSTM-based sequence model) to derive a temporal tension curve. Simultaneously, the rhythmic regularity index of the segment is calculated. These features are then combined to generate the acoustic profile information vector V_audio for the current 2-4 measures. Pipeline B (Score Parsing): Based on the current performance progress, the corresponding measure's score data is extracted from the score information (MusicXML format). Using a rule engine and Natural Language Processing (NLP) technology, tempo terms, dynamic markings, and expression terms are identified and parsed, transforming them into a structured set of emotional intensity and style description instructions, S_score. Subsequently, the system inputs V_audio and S_score together into a pre-built AI music emotion mapping knowledge base. At the core of this knowledge base is a trained multimodal fusion neural network model. It first maps V_audio to an emotional latent space, and simultaneously maps S_score to the same space through an embedding layer. Then, it performs a co-mapping of the two through an attention mechanism, that is, it calculates the consistency weight of audio features and musical scores in emotional expression. Finally, the model performs verification and comprehensive decision-making based on the weighted features, and outputs the most likely emotional context label (e.g., label = {emotion: passionate, intensity: medium to high, confidence: 0.85}).
[0103] In alternative or modified implementations, AI audio analysis techniques can be replaced by using pre-trained large-scale music audio models (such as the encoders of MusicLM or Jukebox) to extract richer semantic features. The parsing of score information can be expanded to include deeper analysis of harmonic progressions and formal structures to obtain more macroscopic emotional cues. The construction of a pre-built AI music emotion mapping knowledge base can be independent of end-to-end neural networks, employing a hybrid system based on knowledge graphs and rule-based reasoning, particularly to incorporate more musicological theory and differences in emotional expression across different cultural backgrounds. The logic of collaborative mapping and verification can incorporate the performer's personal historical data, making emotion inference more personalized.
[0104] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0105] Figure 3 A schematic diagram of an AI-adaptive electric piano touch force calibration system provided in one embodiment of this application is shown below. Figure 3 As shown, an AI-adaptive electric piano touch force calibration system 300 of this embodiment includes: a performance digital fingerprint module 301, a graded calibration strategy module 302, and a calibration suggestion report module 303.
[0106] The performance digital fingerprint module 301 is used to acquire a user's multi-source performance information set, perform identity-seamless collaborative identification and initial state evaluation based on the user's multi-source performance information set, and generate a user performance digital fingerprint information set. The graded calibration strategy module 302 is used to perform multi-dimensional performance ability assessment and dynamic guidance strategy generation based on the user's performance digital fingerprint information set, and generate a personalized graded calibration strategy information set. The calibration suggestion report module 303 is used to generate a personalized touch force calibration suggestion report based on the personalized graded calibration strategy information set, combined with real-time music context perception and emotional intent understanding, to perform forward-looking dynamic force curve adjustment and optimization feedback.
[0107] The system in this embodiment can be used to execute the methods of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
Claims
1. A method for calibrating the touch force of an electric piano based on AI adaptive technology, characterized in that, include: Obtain a multi-source playing information set of the user, and based on the multi-source playing information set of the user, perform identity-insensitive collaborative identification and initial state assessment to generate a digital fingerprint information set of the user's performance. Based on the user's digital fingerprint information set, a multi-dimensional performance ability assessment and dynamic guidance strategy are generated, resulting in a personalized graded calibration strategy information set. Based on the personalized graded calibration strategy information set, combined with real-time music context perception and emotional intent understanding, a forward-looking dynamic dynamic force curve adjustment and optimization feedback is generated, and a personalized touch force calibration suggestion report is generated and output.
2. The method according to claim 1, characterized in that, The process of generating the user's digital fingerprint information set includes: The mobile application client acquires multi-dimensional real-time playing data generated by the user's interaction with the electric piano. Based on the multi-dimensional real-time playing data, an intelligent recognition algorithm is used to determine the user's identity status and construct the user's identity identifier and its preliminary attribute tags. Simultaneously, based on the multi-dimensional real-time playing data, initial key-touch behavior fingerprints are extracted to construct a set of playing features describing the user's basic playing habits and ability tendencies. By combining the user's identity identifier and its preliminary attribute tags with the set of playing features, the user's performance digital fingerprint information set is generated.
3. The method according to claim 2, characterized in that, The process of constructing the user identity identifier and its preliminary attribute tags includes: Based on the historical stored data of the mobile application client, determine whether a verified account session initiated through the associated application currently exists: If it exists, the current session account information will be directly used as the user identity identifier, and historical performance maturity tags will be loaded from the cloud archive as the preliminary attribute tags; If not found, the anonymous perception process is initiated. Through AI feature comparison, the user's initial free-playing segment is analyzed to extract its typical dynamic distribution and rhythmic stability features, and compared with the local temporary archive. If a match is found, the corresponding anonymous identifier and label will be used. If no match is found, a new anonymous identifier is created, and based on the typical intensity distribution and the rhythmic stability characteristics, a preliminary attribute label is constructed to identify its basic proficiency level.
4. The method according to claim 2, characterized in that, The process of constructing the set of playing features includes: From the multidimensional real-time playing data, the basic dynamic distribution range of the user when playing continuously without guidance is extracted, and the frequently used dynamic values and their dispersion are analyzed. Synchronously extract users' typical rhythm control patterns, analyze their time deviation rate when playing even notes, and their basic playing accuracy when handling simple rhythm patterns; The basic dynamic range, the frequently used dynamic values and their dispersion, the time deviation rate, and the basic playing accuracy are correlated and encoded to construct the playing feature set, which is used to characterize the user's original control habits and stability tendency without intervention.
5. The method according to claim 2, characterized in that, The process of generating the personalized graded calibration strategy information set includes: Based on the set of playing features and the preliminary attribute labels, a performance maturity profile containing specific intensity indicators and weak link identifiers is constructed through multi-dimensional performance ability assessment. Based on the performance maturity profile, a set of dynamic calibration rules corresponding to the level is mapped and constructed through an AI-driven three-level guidance strategy generation logic. The performance maturity profile and the set of dynamic calibration rules are encapsulated to generate the personalized graded calibration strategy information set.
6. The method according to claim 5, characterized in that, The process of constructing a performance maturity profile through multi-dimensional performance ability assessment includes: Based on the set of playing features, the user's control accuracy under different preset force targets is analyzed, and a force control stability score is generated. Analyze the degree of match between the actual output dynamic range and the target range when users process musical phrases containing clear contrasts in volume, and generate a dynamic response awareness score; Analyze the clarity and evenness of notes when users attempt basic techniques such as fast scales or arpeggios, and generate a basic technique performance score; By combining the strength control stability score, the dynamic response awareness score, and the basic skill execution score, along with the preliminary attribute labels, a performance maturity profile represented by a radar chart is constructed.
7. The method according to claim 6, characterized in that, The AI-driven three-level guidance strategy generation logic constructs a dynamic calibration rule set, including: If the performance maturity profile indicates that the user is in the stage of forming basic control ability, then auxiliary shaping strategy rules are generated. The core rules include: identifying the continuous too light key touch range and generating visual feedback suggestions to prompt the user to increase the pressing pressure, and identifying the too heavy key touch that is easy to cause tone breakage and generating auditory examples to prompt the user to control the pressure. If the performance maturity profile indicates that the user is in the stage of expanding control ability, then a collaborative reinforcement strategy rule is generated. The core rules include: when the user practices strong phrases, marking the difference between the peak dynamics and the target peak dynamics and providing dynamic range practice prompts; when the user practices weak phrases, highlighting the dynamic fluctuation range and providing stability control suggestions. If the performance maturity profile indicates that the user is in a stable and refined control stage, then a transparent resonance strategy rule is generated. The core rule includes: recording the user's distinctive dynamic techniques in the performance and comparing and analyzing them with standard interpretations to generate a descriptive feedback report for improving the refinement of artistic expression.
8. The method according to claim 7, characterized in that, The process of generating and outputting the personalized key pressure calibration suggestion report includes: Based on the audio stream and score information of real-time performance, emotional context labels describing the emotional tone of the current musical passage are constructed through real-time musical context perception. Based on the emotional context tags and the set of dynamic calibration rules, instantaneous calibration parameters that adapt to the current musical emotion are generated through AI adaptive dynamic curve tuning logic. Based on the comparative analysis of the actual dynamics and the expected dynamics of the emotional context labels throughout the entire performance segment, a text report containing specific improvement ranges and descriptive suggestions is generated, and the change history of the instantaneous calibration parameters is attached to the report. The text report and its attachments are visualized and pushed through the mobile application client to generate and output the personalized touch force calibration suggestion report.
9. The method according to claim 8, characterized in that, The process of constructing the emotional context tags includes: AI audio analysis technology is used to analyze the audio stream of the real-time performance, extract the quantitative features of its chord tension and rhythm regularity, and generate sound profile information that represents the musical atmosphere. The musical score information of the current performance section is analyzed simultaneously, and the tempo terms, dynamic markings and expression terms are identified and converted into clear instructions for emotional intensity and style description. By using a pre-built AI music emotion mapping knowledge base, the sound profile information is collaboratively mapped and verified with the emotion intensity and the style description instructions to generate the emotion context label used to guide subsequent intensity calibration.
10. An AI-adaptive digital piano touch force calibration system, characterized in that, The method applied to any one of claims 1-9 includes: The performance digital fingerprint module is used to acquire a user's multi-source performance information set, and based on the user's multi-source performance information set, to perform identity-insensitive collaborative identification and initial state evaluation, and generate a user performance digital fingerprint information set. The graded calibration strategy module is used to perform multi-dimensional performance ability assessment and dynamic guidance strategy generation based on the user's performance digital fingerprint information set, and generate a personalized graded calibration strategy information set. The calibration suggestion report module is used to generate a personalized touch force calibration suggestion report based on the personalized graded calibration strategy information set, combined with real-time music context perception and emotional intent understanding, to perform forward-looking dynamic force curve adjustment and optimization feedback.