Piano music expression teaching method and system based on emotion calculation

By constructing a music semantic association network and aligning semantic features in a cross-modal joint semantic space, the problem that existing systems cannot convert the emotional mood of music into performance control means is solved, realizing intelligent and precise guidance for piano teaching and improving the scientific nature and effectiveness of teaching.

CN121858962APending Publication Date: 2026-04-14CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing music-assisted teaching systems cannot effectively establish a definite logical mapping relationship between high-level semantic understanding and low-level physical characteristics, which makes it impossible to convert the emotional mood in the background of music into specific performance control means, thus limiting the depth and breadth of application of intelligent music teaching systems in the level of emotional expression.

Method used

By using an emotion-based computing approach, a music semantic association network is constructed to generate structured semantic feature vectors and acoustic physical feature vectors. Logical alignment is achieved within a cross-modal joint semantic space, and quantitative performance control parameters, including key touch force values ​​and note duration offsets, are output.

Benefits of technology

It achieves a precise mapping from implicit emotional imagery to explicit performance techniques, provides personalized and operationally oriented teaching guidance, and shortens the learning cycle from technical practice to artistic expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858962A_ABST
    Figure CN121858962A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer-aided teaching, discloses a piano music expression teaching method and system based on emotion calculation, remarkably improves scientificity and effectiveness of piano music emotion expression teaching, effectively solves the problem that art intentions are difficult to quantitatively teach in traditional teaching, and improves the teaching efficiency. By constructing a music semantic association network fused with music theory logic, the system breaks through a semantic gap between an abstract music text and a specific acoustic signal, realizes accurate mapping from a recessive emotional artistic conception to a dominant playing technique, and utilizes a cross-modal joint semantic space and a reverse parameter analysis technology to realize accurate mapping of the music semantic association network. According to an emotion background of a target track, a high-dimensional artistic concept can be automatically decoupled into executable physical control parameters such as key touch force and time value deviation, and personalized guidance with clear operation directivity is provided for students.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer-aided instruction technology, and in particular to a piano music expression teaching method and system based on affective computing. Background Technology

[0002] The art of piano performance not only demands precise fingering and rhythmic control, but also relies on a profound understanding and artistic representation of the music's deeper emotions and the composer's intentions. In traditional higher music education and professional piano teaching, teachers often inspire students' artistic imagination by explaining the background of the work's composition, analyzing its form, and using rich literary descriptions. For example, teachers might explain the grand historical context of Beethoven's works, filled with contradictions and conflicts, guiding students to understand the spiritual core behind the score and thus infuse their performance with corresponding emotional tension. Simultaneously, with the rapid development of digital music technology, a vast amount of recordings of renowned performances and detailed musicological literature have been digitized, forming a massive multimodal music data resource. This data contains a complete knowledge system, ranging from macroscopic descriptions of emotional atmosphere to microscopic handling of actual performance techniques. Theoretically, it can provide standardized artistic expression references for students lacking face-to-face guidance from renowned teachers, helping them transition from purely mechanical technical training to higher-level artistic expression. This is a key foundation for the intelligent transformation of modern music education.

[0003] However, existing music-assisted teaching and information processing technologies face insurmountable technical obstacles in unlocking the educational value of these cross-modal data. Current mainstream solutions typically treat textual and audio data as two separate domains, processing them in isolation. On the one hand, natural language processing (NLP) is used to classify music documents based on keyword-based sentiment, while on the other hand, audio signal processing is used to extract physical acoustic features such as frequency and amplitude from recordings. This fragmented approach leads to a severe disconnect between high-level semantic understanding and low-level physical features, preventing the system from establishing a definite logical mapping between abstract literary descriptions and specific acoustic parameters of performance. Specifically, while existing NLP algorithms can identify emotional tendencies in text, they lack the ability to deeply analyze musical terminology and metaphorical rhetoric. They cannot accurately quantify high-level semantic concepts such as heroism, inner tension, or sighing phrases into computationally computable and executable physical parameter constraints such as keystroke speed, micro-dynamic curves, and the range of free rhythmic control. Due to the lack of such cross-modal fine-grained semantic alignment mechanisms, computers struggle to parse out the specific performance control methods needed to express the emotional mood within a particular context. This results in the teaching system being unable to generate performance suggestions with substantial operational guidance based on the musical background, thus limiting the depth and breadth of application of intelligent music teaching systems in the realm of emotional expression. Summary of the Invention

[0004] This application proposes a piano music expression teaching method and system based on affective computing to address the problems raised in the background art.

[0005] To achieve the above objectives, this application adopts the following technical solution: a piano music expression teaching method based on affective computing, comprising the following steps:

[0006] Step S1: Obtain multimodal music data containing musicological text data and demonstration audio data; perform entity extraction and relation parsing on the musicological text data; establish logical connections between emotion description entities and performance technique entities; and construct a music semantic association network.

[0007] Step S2: Using the music semantic association network as constraint prior information, feature encoding is performed on the musicological text data to generate a structured semantic feature vector containing topological structure information. At the same time, acoustic signal analysis is performed on the demonstration audio data to generate an acoustic physical feature vector that matches the dimension of the structured semantic feature vector.

[0008] Step S3: Calculate the semantic similarity between the structured semantic feature vector and the acoustic physical feature vector. By minimizing the distance between semantically corresponding samples and maximizing the distance between semantically non-corresponding samples, a cross-modal joint semantic space is established to achieve logical alignment between text semantics and acoustic features in this cross-modal joint semantic space.

[0009] Step S4: In response to the emotion expression query command for the target piece, retrieve the target acoustic feature region that matches the vector representation of the emotion expression query command in the cross-modal joint semantic space, and perform inverse parameter parsing on the target acoustic feature region to output quantized performance control parameters, which include the touch force value and the note duration offset.

[0010] Furthermore, in step S1, the steps of entity extraction and relation parsing of the musicological text data specifically include:

[0011] First, a pre-set high-dimensional music ontology vector basis space is loaded, and text fragments in the musicology text data are mapped to the high-dimensional music ontology vector basis space. The musicology information density score is determined by calculating the projection distance between the text semantic vector and the music ontology basis vector. The musicology information density score is compared with a preset density threshold, and text fragments with scores lower than the density threshold are removed to obtain high-density text data.

[0012] Subsequently, a hierarchical multi-task extraction architecture is run for high-density text data. Entity categories are pre-divided into macro-sentiment level, meso-structure level, and micro-technique level. During the extraction process, a hierarchical dependency mask constraint mechanism is introduced to calculate the conditional activation probability of an entity belonging to a specific category. This forces the recognition of micro-technique level entities to depend on the contextual information of macro-sentiment level entities, thereby outputting a set of candidate entity nodes with hierarchical attribute labels.

[0013] Furthermore, in step S1, the step of establishing a logical connection between the emotion description entity and the performance technique entity, and constructing a music semantic association network, specifically includes:

[0014] The entity pairs in the candidate entity node set are traversed. Natural language dependency parsing is used to identify the modification and modified relationships between entity words. At the same time, a music theory logic verification operator is introduced to query the preset music theory rule base to verify the acoustic rationality of the relationship between emotion description entities and performance technique entities. The confidence of logical relationship edges is calculated by fusing syntactic dependency scores and music theory logic consistency scores. Only logical relationship edges that pass the verification are retained.

[0015] Subsequently, based on the retained logical relationship edges, a global topological energy minimization pruning operation is performed to construct a global semantic consistency potential energy function that reflects the overall stability of the network. The semantic centrality index of each entity node is calculated, and weak connection subgraphs that conflict with the global sentiment theme are automatically removed. The remaining entity nodes and logical relationship edges after pruning are assembled into a structurally stable music semantic association network.

[0016] Furthermore, in step S2, the step of using the music semantic association network as constraint prior information to perform feature encoding on the musicological text data and generate a structured semantic feature vector containing topological structure information specifically includes:

[0017] A graph-guided semantic encoder is constructed to map words in musicological text data to initial embedding vectors and align them to the corresponding entity nodes in the music semantic association network. Then, a topological semantic diffusion aggregation operation is performed to extract the confidence of logical relationships corresponding to logical connections in the music semantic association network as transit weights. The semantic information of the current entity node in the network neighborhood is aggregated using a graph convolution mechanism.

[0018] In the topological semantic diffusion aggregation operation, the semantic features of the nodes themselves are extracted using the self-loop feature transformation matrix, and the topological features diffused from the neighboring nodes are extracted using the neighborhood feature transformation matrix. The topological injection strength coefficient is introduced to numerically adjust the degree of correction of the network structure information. The fused features are processed by a nonlinear activation function to generate a structured semantic feature vector containing rigorous topological logic.

[0019] Furthermore, in step S2, the step of performing acoustic signal analysis on the demonstration audio data to generate acoustic physical feature vectors that are adapted to the dimensions of the structured semantic feature vectors specifically includes:

[0020] A cross-modal technique prior distillation mechanism is adopted to retrieve semantic embedding vectors belonging to performance technique entities from the music semantic association network, and calculate the aggregate mean of semantic embedding vectors to generate global technique prior vectors.

[0021] Subsequently, the global technique prior vector is used as a gating signal to constrain the acoustic analysis process of the demonstration audio data. Specifically, the dot product correlation operation is performed between the Mel spectrogram features of the demonstration audio data and the global technique prior vector after cross-modal projection to generate a heatmap of attention indicating key frequency band regions. The heatmap of attention is then used to weight and mask the original Mel spectrogram features, forcing the convolutional neural network to focus on the time-frequency region that is strongly related to the performance technique. Finally, the fully connected layer maps and outputs an acoustic physical feature vector whose dimension is strictly consistent with the structured semantic feature vector.

[0022] Furthermore, in step S3, the step of calculating the semantic similarity between the structured semantic feature vector and the acoustic physical feature vector specifically includes:

[0023] First, a joint manifold space projection operation is performed, introducing the text manifold projection matrix and the acoustic manifold projection matrix as nonlinear mapping heads, which respectively map the structured semantic feature vector and the acoustic physical feature vector onto the shared unit hypersphere.

[0024] Subsequently, a semantic acoustic consistency metric function is defined and computed. During the computation, a bilinear interaction tensor is introduced to capture the second-order interaction relationship between each dimension of the feature vector. The semantic similarity score, which indicates the degree of sample matching, is obtained by calculating the bilinear product between the projected structured semantic feature vector and the projected acoustic physical feature vector, and dividing the bilinear product by the modulus product adjusted by the temperature scaling factor.

[0025] Furthermore, in step S3, the specific steps of establishing a cross-modal joint semantic space by minimizing the distance between semantically corresponding samples and maximizing the distance between semantically non-corresponding samples include:

[0026] Perform spatial alignment operations and construct a topology-adaptive contrastive loss function. Utilize the music semantic association network to calculate the graph geodesic distance between any two entity nodes and transform the graph geodesic distance into semantic neighborhood weights.

[0027] In the process of calculating the topological adaptive contrastive loss function, the semantic similarity score of semantically corresponding positive sample pairs is maximized, and for semantically non-corresponding negative sample pairs, an exponential function containing a semantic decay feature constant is constructed to construct a topological repulsion factor. The repulsion strength of negative sample pairs in the vector space is dynamically adjusted according to the graph geodesic distance, so that negative sample pairs that are closer in distance in the music semantic association network are subjected to less repulsion. Finally, by minimizing the total topological adaptive contrastive loss, the vector distribution in the joint space is forced to conform to the topological logic of the music semantic association network, thereby establishing a cross-modal joint semantic space.

[0028] Furthermore, in step S4, the step of retrieving the target acoustic feature region in the cross-modal joint semantic space that matches the vector representation of the emotion expression query instruction in response to the emotion expression query instruction for the target track specifically includes:

[0029] A graph-guided semantic encoder is reused to convert the received emotional expression query instructions for the target track into a high-dimensional query intent vector;

[0030] Subsequently, a semantic resonance region scanning operation is performed in the cross-modal joint semantic space to calculate the semantic resonance degree between the query intent vector and each pre-stored acoustic physical feature vector in the space. The multi-head attention mechanism is used to weight and aggregate the acoustic physical feature vectors that present high resonance degree, thereby locking in the target acoustic feature region that best represents the emotional expression query instruction.

[0031] Furthermore, in step S4, the specific steps of performing inverse parameter analysis on the target acoustic feature region and outputting quantized performance control parameters include:

[0032] An adaptive gating inverse decoupling network is constructed, and the query intent vector is used as the driving signal input to the gating weight matrix. After processing by the hyperbolic tangent activation function, a semantic feature filtering gating for a specific performance dimension is generated.

[0033] Using semantic feature filtering gates, element-wise multiplication operations are performed on the target acoustic feature region to extract latent feature components that are strongly related to performance control from the abstract acoustic features. The latent feature components are then input into the inverse decoding matrix. The high-dimensional feature space is mapped back to the low-dimensional control parameter space through matrix projection transformation, and the instrument's inherent deviation vector is superimposed.

[0034] Subsequently, a physical constraint transformation operator is introduced to post-process the mapped values, forcing the dimensionless neural network output values ​​to be mapped into the physically feasible domain of piano performance biomechanics, and finally generating quantized performance control parameters. These quantized performance control parameters are specifically represented as a time series matrix containing two independent channels, where the first channel is the touch force value conforming to the instrument digital interface standard, and the second channel is the note value offset in milliseconds.

[0035] The piano music expression teaching system based on affective computing specifically includes: a music semantic association network construction module, a cross-modal feature encoding and generation module, a cross-modal joint semantic space construction module, and an affective expression retrieval and inverse parameter parsing module, wherein;

[0036] The music semantic association network construction module is used to acquire multimodal music data containing musicological text data and demonstration audio data, perform entity extraction and relation parsing on the musicological text data, establish logical connections between emotional description entities and performance technique entities, and construct a music semantic association network.

[0037] The cross-modal feature encoding and generation module is used to encode the music semantic association network as constrained prior information, generate a structured semantic feature vector containing topological information, and perform acoustic signal analysis on the demonstration audio data to generate an acoustic physical feature vector that matches the dimension of the structured semantic feature vector.

[0038] The cross-modal joint semantic space construction module is used to calculate the semantic similarity between structured semantic feature vectors and acoustic physical feature vectors. By minimizing the distance between semantically corresponding samples and maximizing the distance between semantically non-corresponding samples, a cross-modal joint semantic space is established to achieve logical alignment between text semantics and acoustic features within this cross-modal joint semantic space.

[0039] The emotion expression retrieval and inverse parameter parsing module is used to respond to the emotion expression query command for the target song, retrieve the target acoustic feature region that matches the vector representation of the emotion expression query command in the cross-modal joint semantic space, perform inverse parameter parsing on the target acoustic feature region, and output quantized performance control parameters, which include the touch force value and note duration offset.

[0040] The beneficial effects of this invention are as follows:

[0041] This invention significantly enhances the scientific rigor and effectiveness of teaching emotional expression in piano music, effectively addressing the pain point of traditional teaching where artistic intent is difficult to quantify and impart. By constructing a musical semantic association network that integrates music theory logic, the system breaks down the semantic gap between abstract musicological texts and specific acoustic signals, achieving a precise mapping from implicit emotional mood to explicit performance techniques. Utilizing cross-modal joint semantic space and inverse parameter analysis technology, the system can automatically decouple high-dimensional artistic concepts into executable physical control parameters such as touch intensity and timing offset based on the emotional background of the target piece. This provides students with personalized guidance with clear operational direction. This mechanism not only ensures the rigor of teaching suggestions in musical logic but also allows students to intuitively understand the touch control principles behind emotional expression, thereby significantly shortening the learning cycle from technical practice to artistic expression. This achieves a leapfrog development in piano teaching from empiricism to intelligent and precise teaching. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort:

[0043] Figure 1 This is a flowchart of the method of the present invention;

[0044] Figure 2 This is a logic diagram of step one of the present invention;

[0045] Figure 3 This is the logic diagram for step two of the present invention;

[0046] Figure 4 This is the logic diagram for step three of the present invention;

[0047] Figure 5 This is the logic diagram for step four of the present invention;

[0048] Figure 6 This is a system framework diagram of the steps of the present invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] Example 1, as Figure 1 As shown, the piano music expression teaching method based on affective computing includes the following steps:

[0051] Step S1: Obtain multimodal music data containing musicological text data and demonstration audio data; perform entity extraction and relation parsing on the musicological text data; establish logical connections between emotion description entities and performance technique entities; and construct a music semantic association network.

[0052] Step S2: Using the music semantic association network as constraint prior information, feature encoding is performed on the musicological text data to generate a structured semantic feature vector containing topological structure information. At the same time, acoustic signal analysis is performed on the demonstration audio data to generate an acoustic physical feature vector that matches the dimension of the structured semantic feature vector.

[0053] Step S3: Calculate the semantic similarity between the structured semantic feature vector and the acoustic physical feature vector. By minimizing the distance between semantically corresponding samples and maximizing the distance between semantically non-corresponding samples, a cross-modal joint semantic space is established to achieve logical alignment between text semantics and acoustic features in this cross-modal joint semantic space.

[0054] Step S4: In response to the emotion expression query command for the target piece, retrieve the target acoustic feature region that matches the vector representation of the emotion expression query command in the cross-modal joint semantic space, and perform inverse parameter parsing on the target acoustic feature region to output quantized performance control parameters, which include the touch force value and the note duration offset.

[0055] Example 2, as Figure 2 As shown, step S1, which involves entity extraction and relation parsing of the musicological text data, specifically includes:

[0056] The system first loads a pre-defined high-dimensional music ontology vector basis space, maps text fragments from the musicological text data to this space, and determines the musicological information density score by calculating the projection distance between the text semantic vectors and the music ontology basis vectors. In this process, to accurately quantify the "value" of the text, the system introduces the following calculation formula:

[0057] ;

[0058] In this formula:

[0059] Representing the The musicological information density score of each input text fragment, with a value range of [value missing]. .

[0060] : Represents the current text segment The total number of valid words contained in the result is used to normalize the length of the accumulated result, eliminating the impact of text length differences on the score.

[0061] Represents words in the text A high-dimensional semantic feature vector. In this embodiment, this vector is generated by a pre-trained deep language model (such as BERT or Word2Vec), and its dimension is... The preferred values ​​are 768 or 1024.

[0062] This represents the pre-defined normal vector of the music ontology hyperplane. This vector is a unit vector pre-calculated based on a professional music literature corpus (such as the *New Grove Dictionary of Music and Musicians*) using principal component analysis (PCA) or cluster center extraction algorithms. Its direction represents the optimal projection direction of "musicological knowledge" in the high-dimensional semantic space. In the formula... Item is the word to be counted The semantic projection strength in this professional direction.

[0063] : Represents the scaling factor, i.e., the vector dimension. The square root of the result. The purpose of introducing this factor is to prevent the sigmoid activation function from entering the saturation region (gradient vanishing) due to excessively large values ​​of the dot product of high-dimensional vectors, thereby ensuring the sensitivity of the scoring.

[0064] : Representing words Normalized term frequencies in a pre-defined corpus of professional music terminology. Logarithmic terms are introduced. This is to conform to Zipf's Law in linguistics, smoothing out the order-of-magnitude difference between high-frequency and low-frequency words.

[0065] : Represents the semantic-statistical balance coefficient (adjustment factor), used to adjust the weight ratio of semantic projection features and statistical word frequency features in the total score. In a preferred embodiment of the present invention, The value range is set to 0.3 to 0.8, with an optimal value of 0.5. When Setting it to this range effectively balances semantic relevance and the frequency of use of technical terms, avoiding interference from common high-frequency words while preventing obscure technical terms from being missed.

[0066] Sigmoid: Represents the logistic activation function, used to map the result of the linear combination within the parentheses to... Interval, as a word The independent density contribution value.

[0067] Next, the musicological information density will be scored. With the preset density threshold Numerical comparisons were performed, and those with scores below the density threshold were removed. This step obtains high-density text data by extracting text fragments; this step ensures that subsequent processing is only performed on high-value text rich in emotional adjectives and technical terms, effectively removing noise interference.

[0068] In summary, this formula, by integrating semantic projection in geometric space with word frequency distribution in statistical space, achieves accurate calculation of the "musicality content" of a text. The system ultimately calculates the... With the preset density threshold (In this embodiment) The preferred setting is 0.65) for comparison. If the data is high-density, it is considered high-density text data and retained; otherwise, it is considered noise data and discarded.

[0069] Subsequently, a hierarchical multi-task extraction architecture was implemented for high-density text data. Entity categories were pre-divided into macro-level sentiment (L1), meso-level structure (L2), and micro-level technique (L3). During the extraction process, a hierarchical dependency mask constraint mechanism was introduced to calculate the conditional activation probability of an entity belonging to a specific category. To achieve this hierarchical constraint, this implementation uses a BERT-based sequence labeling model and applies the following probability calculation formula:

[0070] ;

[0071] In this formula:

[0072] : Represents the context of a given global environment In the case of the current entity to be identified Classified into a specific category The conditional probability. This probability value indicates the confidence level of the entity classification.

[0073] : Represents the category The learnable weight matrix. This matrix is ​​obtained through model training and is used to map the feature vectors of entities to the category space.

[0074] : Represents the entity to be identified The contextual embedding vector. In this embodiment, the vector is output by the encoding layer of a pre-trained language model such as BERT, and the preferred dimension is 768. It contains the semantic information of the entity itself.

[0075] : Represents the Hierarchical Constraint Coefficient. This is a key hyperparameter used to adjust the weight of the influence of parent-level context information on the classification decision of the current child level. In this embodiment, The preferred value range is 0.8 to 1.2, with an optimal value of 1.0. The purpose of introducing this coefficient is to strengthen the constraint of the macro-emotional context and prevent the model from incorrectly identifying incompatible micro-techniques in the absence of corresponding emotional tone support. For example, it prevents the incorrect identification of the "loud voice" technique entity in a "calm" macro-context.

[0076] : Represents the current target category The set of parent hierarchy categories. For example, if the category The "Micro Techniques Layer (L3)" is the set of its parent layers. This includes the "macro-sentiment layer (L1)". This defines the dependency paths in the knowledge graph.

[0077] : Represents the parent node attention weight. When multiple parent-level cues exist, this weight is used to measure the contribution of different parent-level entities to the current child-level decision. This parameter is automatically learned and updated during model training and satisfies the following conditions: .

[0078] : Represents the feature representation vector of the parent-level entity in the context. The existence of this term ensures that the classification calculation of the child level must explicitly aggregate the semantic features of the parent level, thereby achieving logical decoupling and alignment of "emotion and technique".

[0079] Softmax: Represents the normalization exponential function, used to convert the result of linear operations within parentheses into a probability distribution form.

[0080] In summary, this formula explicitly adds a function derived from a conventional classifier to the existing classifier. The weighted parent-level contextual terms force the model to learn the prior logic of "emotion determining technique," thereby greatly improving the accuracy and logical consistency of entity extraction. In this implementation, the formula forces the recognition of micro-level technique entities to depend on the contextual information of macro-level emotion entities, thus outputting a set of candidate entity nodes with hierarchical attribute labels. Through this Bayesian prior constraint, the system can avoid ambiguity; for example, "strong voice" is only marked as a valid entity when a macro-emotional context of "excitement" is detected.

[0081] Furthermore, in step S1, a logical connection is established between the emotion description entity and the performance technique entity to construct a music semantic association network. The specific steps include:

[0082] Traverse the set of candidate entity nodes The system uses natural language dependency parsing to identify the modification and being modified relationships between entity pairs. Simultaneously, it incorporates a music theory logic verification operator to query a pre-defined music theory rule base to verify the acoustic rationality of the association between emotion description entities and performance technique entities. In this stage, the system calculates the confidence level of logical relationship edges by fusing syntactic dependency scores and music theory logic consistency scores. The specific calculation model is as follows:

[0083] ;

[0084] In this formula:

[0085] : Represents an entity and The final confidence probability of a logical relationship exists, and its value ranges from [value missing]. The system only retains relation edges whose value is greater than a preset threshold (e.g., 0.75).

[0086] This represents the syntactic dependency score. This value is generated by the dependency parser in the natural language processing module and indicates the syntactic dependency score. Modification The probability of, for example, the dependency probability of “sad (adjective)” modifying “melody (noun)”.

[0087] : Represents a multimodal fusion operator. In this embodiment, algebraic multiplication is preferably used. The multiplication mechanism ensures that the final confidence level is high only when both the "grammatical structure is reasonable" and the "musical logic is consistent" conditions are met simultaneously, thus achieving strict logical filtering.

[0088] : Represents a music theory attribute mapping function. This function maps text entities to a predefined "music physical attribute space." This space typically contains three core dimensions: "Energy," "Tension," and "Valence." For example, the vector mapped from the entity "excitement" The values ​​in the energy dimension are higher, while the mapping vector of the entity "flexible plate" has lower values ​​in the energy dimension.

[0089] : Represents the square of the Mahalanobis distance. This term is used to calculate the degree of difference between two entities in the music-physical property space. This is the covariance matrix of the attribute space, used to eliminate the influence of inconsistent dimensions between different attribute dimensions (such as velocity and intensity). The larger this distance value, the more conflicting the two entities are in terms of music theory, such as the conflict between "tranquility" and "loudness".

[0090] : Represented by the natural constant An exponential function with base 1 converts the above distance values ​​into... The similarity probability of intervals. The smaller the distance (the smaller the contradiction), the closer this term is to 1.

[0091] : Represents the music theory tolerance coefficient (i.e., the variance of the Gaussian kernel function). This parameter controls the system's sensitivity to music theory conflicts. In this embodiment, The preferred setting is 0.5 to 1.0. The smaller the value, the stricter the system's requirements for music theory logic, and the more effectively it can filter out literary descriptions that are rhetorically exaggerated but do not conform to acoustic laws.

[0092] In summary, this formula, by fusing syntactic probabilities from the NLP domain with logistic probabilities from the music acoustics domain, achieves dual verification of edges in the knowledge graph, ensuring that the constructed music semantic association network is logically rigorous and executable. This formula quantifies the consistency between linguistic associations and musicological logic; the system then retains only the confidence score based on this. By verifying the logical relationship edges, we can eliminate rhetorical fallacies commonly found in literary descriptions and ensure that the preserved relationships have executable musical significance.

[0093] Subsequently, a global topological energy minimization pruning operation is performed based on the preserved logical relationship edges to construct a global semantic consistency potential function that reflects the overall stability of the network. The specific definition is as follows:

[0094] ;

[0095] In this formula:

[0096] : Represents the overall state energy of the entire music semantic association network. The optimization objective of this invention is to iteratively adjust the node embedding vectors. By pruning high-energy edges, the potential energy value is minimized, thereby enabling the network to reach topological steady state.

[0097] : Represents the set of all logical relationship edges and the set of all entity nodes in the current network, respectively.

[0098] : Represents the connecting entity and entity The weight of the edge. This weight value is directly inherited from the logical relation confidence score calculated in the previous step. Weight The larger the value, the closer the two nodes are in semantic logic, and the higher their contribution to the distance constraint.

[0099] : Represents a node and The squared Euclidean distance in the final high-dimensional semantic space. The first term of the formula (i.e., the Laplace regularization term) forces logically strong correlations (i.e., Large nodes are clustered in the vector space to ensure the smoothness of local semantics.

[0100] : Represents the regularization coefficient. This parameter is used to balance the weights between the network's internal smoothness (the first term) and the consistency of the prior distribution (the second term). In this embodiment, The preferred value range is 0.01 to 0.1, with the optimal value being 0.05.

[0101] : Represents relative entropy (Kullback-Leibler Divergence). This function is used to quantify the difference between two probability distributions.

[0102] : Represents the current node Posterior semantic distribution during network optimization.

[0103] : Represents the prior distribution based on a general music knowledge base (such as axioms of music theory). The second term in the formula serves to prevent the network from overfitting the current specific text data and to ensure that the generated semantic structure does not violate general musicological common sense. For example, it prevents the probability distribution of "extremely fast speed" and "sadness" from being forcibly overlapped in the absence of strong evidence.

[0104] The system minimizes this global semantic consistency potential function. The semantic centrality index of each entity node is calculated, weakly connected subgraphs that conflict with the global sentiment theme are automatically removed, and the remaining entity nodes and logical relationship edges after pruning are assembled into a structurally stable music semantic association network. This process ensures that the final generated network is a highly structured and consistent logical system, providing a solid topological foundation for subsequent cross-modal feature alignment.

[0105] Example 3, as Figure 3 As shown, in step S2, the music semantic association network is... As constraining prior information, feature encoding is performed on musicological text data to generate structured semantic feature vectors containing topological information. The specific steps include:

[0106] The system constructs a graph-guided semantic encoder that maps words in musicological text data to initial embedding vectors and aligns them to a music semantic association network. On the corresponding entity nodes. Then, a topological semantic diffusion aggregation operation is performed to extract the music semantic association network. Confidence of logical relations corresponding to logical connections in the middle As a transmission weight, the semantic information of the current entity node in the network neighborhood is aggregated using the graph convolution mechanism.

[0107] In the topological semantic diffusion aggregation operation, the self-loop feature transformation matrix is ​​used respectively. Extracting the semantic features of the node itself and transforming the matrix using neighborhood features Extract topological features diffused from neighboring nodes and introduce a topological injection strength coefficient. The degree of modification to the network structure information is numerically adjusted. To accurately implement this process, this implementation adopts the following topological semantic diffusion aggregation equation:

[0108] ;

[0109] In this formula:

[0110] : Represents the first The feature matrix output by the layered graph neural network. In this embodiment, the output of the last layer is taken as the final structured semantic feature vector, and its dimension is... The optimal value is 512 or 768 to suit the expression requirements of high-dimensional semantic spaces.

[0111] : Represents a nonlinear activation function. In this embodiment, GELU (Gaussian Error LinearUnit) is preferred. Compared with ReLU, GELU has better smoothness and gradient propagation ability when dealing with the randomness of high-dimensional semantic spaces.

[0112] : Represents the self-loop feature transformation matrix. This matrix is ​​a learnable parameter used to extract the semantic features of the node itself in the previous layer, ensuring that the node retains its ontological semantic information during the aggregation process.

[0113] : Represents the previous level (the first) The node feature input matrix of the first layer. ), It consists of initial word embedding vectors from musicological text data.

[0114] : Represents the topology injection strength coefficient. This hyperparameter controls the magnitude of the correction of a node's semantics by neighborhood topology information. In this embodiment, The optimal value range is 0.5 to 1.0. A larger value results in a generated vector that more closely reflects the global structural features of the graph. A moderate value can effectively balance ontology semantics and structural context.

[0115] : Represents the degree matrix of the music semantic association network. (In the formula...) The terms constitute a symmetric normalized Laplacian operator, used to eliminate the problem of inconsistent feature numerical scales caused by differences in the degree of different nodes, and to prevent gradient explosion or vanishing.

[0116] : Represents the adjacency matrix of the music semantic association network, which characterizes the physical connection state between entity nodes.

[0117] : Represents the Hadamard Product, which is the element-wise multiplication operation of a matrix.

[0118] : Represents the logical confidence mask matrix. The element values ​​of this matrix strictly correspond to the logical relationship confidence scores calculated in step S1. ).pass In terms of computation, the system implements weighted aggregation based on logical reliability, that is, only logically related edges with high confidence are allowed to pass significant feature information, thereby suppressing the interference of noisy connections.

[0119] : Represents the neighborhood feature transformation matrix. This matrix is ​​a learnable parameter used to perform a linear transformation on the normalized and logically weighted neighborhood aggregated features to extract spatial structure features.

[0120] In summary, this formula introduces a logical confidence mask. With topology adjustment coefficient This improves upon traditional graph convolutional networks, ensuring that the generated feature vectors... It encompasses both the literal meaning of the text and a rigorously validated musicological logical structure.

[0121] In this embodiment, the formula dynamically injects the static network structure constructed in step S1 into the text features, ensuring that feature aggregation only propagates along the high-confidence logical path verified by S1. This allows the fused features to be processed by a non-linear activation function to generate a structured semantic feature vector containing rigorous topological logic. ).

[0122] Simultaneously, in step S2, acoustic signal analysis is performed on the demonstration audio data to generate structured semantic feature vectors. Dimensionally adapted acoustic physical feature vectors The specific steps include:

[0123] Employing a cross-modal prior distillation mechanism, this study examines music semantic association networks. The semantic embedding vectors belonging to the entity of performance technique are retrieved, and the aggregate mean of the semantic embedding vectors is calculated to generate a global technique prior vector. This vector represents the common semantic features of all known performance techniques in the current knowledge base.

[0124] Then, the global technique of prior vectors was used. The acoustic analysis process of the demonstration audio data is constrained by the gating signal, specifically by using the Mel spectrogram characteristics of the demonstration audio data. With the cross-modal projection matrix Projected global technique prior vector Perform dot product correlation calculations to generate a heatmap indicating the attention levels of key frequency band regions. The specific computational logic of this process is defined by the following prior guided spectrum-time gating equation:

[0125] ;

[0126] In this formula:

[0127] : Represents the final output acoustic physical feature vector. After processing by this formula, its dimension is mapped to the structured semantic feature vector. Consistency (preferably 512 or 768 dimensions in this embodiment) is required to facilitate the construction of a joint semantic space in subsequent steps.

[0128] : Represents a deep convolutional feature extraction network. In this embodiment, ResNet-18 or CRNN (convolutional recurrent neural network) is preferably used as the backbone architecture to extract high-order abstract features from the gated weighted spectrogram.

[0129] : The Mel-spectrogram feature tensor representing the input example audio data. The size of this tensor in the frequency dimension is denoted as . The size over time varies with the audio length. In this embodiment, the number of Mel bands... The preferred setting is 128 to cover the full range of piano playing.

[0130] : Represents element-wise multiplication under the broadcast mechanism. This operation applies the attention heatmap generated by Sigmoid to the original spectrogram to achieve a "soft gating" mechanism, that is, to retain the energy of key frequency bands related to performance techniques and suppress background noise or irrelevant harmonic interference.

[0131] : Represents the logistic activation function, used to map the correlation calculation results within the parentheses to Generate a heatmap of attention based on the interval.

[0132] : Represents the cross-modal projection matrix. This matrix is ​​a learnable parameter whose function is to map the technique prior vectors in the text modality to the acoustic feature space, making its dimension similar to the frequency dimension of the spectrogram. This allows for matching, thus enabling cross-modal dot product operations.

[0133] : Represents the global technique prior vector. This vector is an aggregated representation of the semantic embeddings of all micro-technique entities identified in step S1. It serves as a "query condition" to guide the model in finding matching feature patterns in the spectrum.

[0134] : Represents the scaling factor, i.e., the frequency dimension. The square root (e.g.) The purpose of introducing this factor is to numerically scale the dot product result, preventing the dot product value from becoming too large due to excessive dimensionality, which would cause the Sigmoid function to enter the saturation region (gradient vanishing), thus ensuring the stability of model training.

[0135] : Represents a learnable bias term used to adjust the activation threshold of the attention mechanism.

[0136] In summary, this formula generates a dynamic spectral mask by calculating the correlation between spectral features and technique priors, thus forcing the convolutional network to... By focusing only on the time-frequency ranges that can reflect specific performance techniques (such as attack speed and sustained vibrato), the ability of acoustic features to represent musical emotions is greatly enhanced.

[0137] This formula utilizes the attention heatmap to analyze the features of the original Mel spectrum. Weighted masks are applied to force the convolutional neural network to focus on time-frequency regions strongly related to performance techniques, such as the transient changes at the onset of a note or the vibrato fluctuations in the sustain. Finally, the output dimension and structured semantic feature vector are mapped through fully connected layers. Strictly consistent acoustic physical eigenvectors This process ensures that the generated audio features have a clear musical semantic orientation, rather than being a blind extraction of physical parameters.

[0138] Example 4, as Figure 4 As shown, in this embodiment, in step S3, the structured semantic feature vector is calculated. Acoustic physical eigenvectors The specific steps for determining semantic similarity between them include:

[0139] First, perform a joint manifold space projection operation, introducing the text manifold projection matrix. Harmony acoustic manifold projection matrix As a nonlinear mapping head, the structured semantic feature vectors are respectively... Acoustic physical eigenvectors Mapping to a shared unit hypersphere. This operation eliminates the geometrical differences in the heterogeneous feature spaces.

[0140] Subsequently, a semantic acoustic consistency metric function is defined and computed, introducing a bilinear interaction tensor during the computation process. To capture the second-order interaction relationships between each dimension of the feature vectors, the structured semantic feature vectors after projection are calculated. With the projected acoustic physical feature vector The bilinear product between them is then divided by the temperature scaling factor. The adjusted modulus product yields the semantic similarity score, which indicates the degree of sample matching. The specific computational logic of this process is defined by the following bilinear manifold similarity equation:

[0141] ;

[0142] In this formula:

[0143] : Represents the first The first text sample and the second The semantic similarity score of each audio sample in the joint space. This score serves as the direct basis for subsequent loss function optimization; a higher value indicates a higher semantic match between the two samples.

[0144] and : respectively refer to the output of step S2. The structured semantic feature vector and the first Each acoustic physical feature vector.

[0145] and : Represent the text manifold projection matrix and the acoustic manifold projection matrix, respectively. These two matrices are learnable parameters whose function is to map heterogeneous text and audio features to the same latent manifold space. This eliminates the geometric distribution differences between modes.

[0146] : Represents the bilinear interaction tensor ( Unlike traditional dot product operations (which only capture the correlation of corresponding dimensions), this invention introduces this tensor to capture the second-order interaction between each dimension of text features and each dimension of acoustic features. This enables the model to learn the potential correlations in unaligned dimensions, thereby significantly improving the non-linear representation capability of the metric, such as capturing the strong correlation between the sentiment features of a specific dimension of the text vector and the energy features of a specific dimension of the acoustic vector.

[0147] : Represents the temperature scaling parameter. This parameter is used to adjust the sharpness of the similarity distribution. In this embodiment, The preferred value is 0.07. This coefficient is introduced based on contrastive learning theory, where a smaller value indicates better performance. The value can amplify the similarity difference between positive and negative samples, forcing the model to focus on distinguishing hard negative samples during training, thereby preventing gradient vanishing and improving alignment accuracy.

[0148] : Represents the Euclidean norm (magnitude) of a vector, used to perform L2 normalization on feature vectors, ensuring that the similarity measure is only related to the direction of the vector (i.e., semantic content) and not to its magnitude (i.e., signal strength).

[0149] In summary, this formula establishes a high-precision metric by combining manifold projection, bilinear interaction, and temperature scaling mechanisms, enabling differentiable fine-grained similarity comparisons of heterogeneous text and audio vectors within the same mathematical framework.

[0150] Meanwhile, in step S3, the specific steps for establishing a cross-modal joint semantic space by minimizing the distance between semantically corresponding samples and maximizing the distance between semantically non-corresponding samples include:

[0151] Perform spatial alignment operations and construct a topology-adaptive contrastive loss function. The music semantic association network constructed in step S1 Calculate the graphical distance between any two entity nodes. and the distance of the map geodetic line It is converted into semantic neighborhood weights.

[0152] Calculating the topology-adaptive contrastive loss function During the process, the semantic similarity score of the corresponding positive sample pairs is maximized. For semantically mismatched negative sample pairs, a feature constant containing semantic decay is used. The topological repulsion factor is constructed using an exponential function, based on the geodesic distance in the graph. The repulsion strength of negative sample pairs in the vector space is dynamically adjusted. The specific calculation model is shown below:

[0153] ;

[0154] In this formula:

[0155] : Represents the total loss for topology-adaptive comparison. The optimization objective of this step is to minimize this loss value, thereby forcing the vector distribution in the joint space to conform to the topological logic of the music semantic association network.

[0156] : Represents the total number of sample pairs in the current model training batch.

[0157] : Represents a positive sample pair (i.e., the semantically corresponding first sample pair) The text and the first The semantic similarity score of each audio sample is calculated using the aforementioned bilinear manifold similarity equation. The numerator of the equation maximizes the closeness of positive sample pairs.

[0158] : Represents a negative sample pair (i.e., the first sample pair that does not correspond semantically). The text and the first The semantic similarity score of each audio file.

[0159] : Represents the entity node in the music semantic association network constructed in step S1. With entity nodes The graph geodesic distance between two samples quantifies the degree of semantic difference between them in real musicological logic. For example, the geodesic distance between "sad" and "desolate" is small, while the geodesic distance between "sad" and "joyful" is large.

[0160] : Represents the semantic decay constant. This parameter controls the sensitivity of the topological repulsion factor to changes in semantic distance. In this embodiment, The preferred setting is 2.0. When At that time, the rejection factor can distinguish negative samples at different distances with a smooth non-linear curve.

[0161] Topological repulsion factor: i.e., in the formula This factor dynamically adjusts the repulsion strength of negative sample pairs in the vector space.

[0162] Among them, when negative samples With sample Distance in the map When the sample is close, the factor approaches 0, thus reducing the repulsive force on the negative sample and allowing semantically similar samples to remain adjacent in the vector space (achieving "soft alignment"); when the distance... When the value is large, the factor approaches 1, thus exerting the greatest repulsive force and pushing away semantically conflicting samples.

[0163] The system achieves logical alignment of cross-modal features by minimizing the total loss of topological adaptive contrast, ensuring that the final cross-modal joint semantic space not only matches numerically but also faithfully reproduces the semantic relationships of musicology in terms of topological structure.

[0164] This formula enables the use of music semantic association networks. Negative sample pairs that are closer in distance experience less repulsion, meaning that "sadness" and "desolation" are allowed to remain close in the vector space, while negative sample pairs that are farther apart experience greater repulsion. Ultimately, the total contrast loss is minimized by minimizing the topological adaptive contrast loss. This forces the vector distribution within the joint space to conform to the semantic association network of music. The topological logic is used to establish a structurally rigorous and logically self-consistent cross-modal joint semantic space.

[0165] Example 5, as Figure 5 As shown, in step S4, in response to the emotion expression query instruction for the target track, a target acoustic feature region matching the vector representation of the emotion expression query instruction is retrieved in the cross-modal joint semantic space. The specific steps include:

[0166] The system first reuses the graph-guided semantic encoder built in step S2 to convert the received emotional expression query instructions for the target track (e.g., "a more explosive heroic tone is needed here") into a high-dimensional query intent vector. .

[0167] Subsequently, a semantic resonance region scan operation is performed in the cross-modal joint semantic space to compute the query intent vector. By analyzing the semantic resonance between the acoustic physical feature vectors and pre-stored acoustic physical feature vectors in space, a multi-head attention mechanism is used to weighted aggregate acoustic physical feature vectors exhibiting high resonance, thereby identifying the target acoustic feature region that best represents the emotional expression query command. This region is not a feature of a single audio segment, but rather an "ideal acoustic prototype" formed by a weighted fusion of features from multiple performance segments by renowned artists that conform to this intent.

[0168] Furthermore, in step S4, the target acoustic feature region is... Perform inverse parameter parsing and output quantized performance control parameters. The specific steps include:

[0169] Construct an adaptive gated inverse decoupling network and utilize query intent vectors. As a driving signal input to the gate weight matrix After passing through the hyperbolic tangent activation function After processing, semantic feature filtering gates are generated for specific performance dimensions.

[0170] Using semantic features to filter gating target acoustic feature regions Element-wise multiplication (Hadamard product) is performed to extract latent feature components strongly correlated with performance control from the abstract acoustic features, and these latent feature components are then input into the inverse decoding matrix. The high-dimensional feature space is mapped back to the low-dimensional control parameter space through matrix projection transformation, and the instrument's inherent deviation vector is superimposed. .

[0171] Subsequently, a physical constraint transformation operator was introduced. The mapped values ​​are post-processed to force the dimensionless neural network output values ​​to be mapped into the physically feasible domain of piano playing biomechanics, such as ensuring non-negativity of dynamics and that timing offsets do not exceed limits. The specific computational logic of this process is defined by the following adaptive semantic-acoustic inverse mapping equation:

[0172] ;

[0173] In this formula:

[0174] : Represents the quantized performance control parameters of the final output. This parameter is a time-series matrix containing two independent channels: Channel 1 (Touch Velocity): Corresponds to the Velocity value in the Musical Instrument Digital Interface (MIDI) standard. Channel 2 (Temporal Offset): Corresponds to the microscopic offset of the note's start time relative to the standard beat. The unit is milliseconds (ms).

[0175] : Represents the physical constraint transformation operator ( The function of this operator is to force the dimensionless numerical values ​​output by the neural network to be mapped to the physically feasible domain of piano playing biomechanics. In this embodiment, the operator includes the following constraint logic:

[0176] Force cutoff: Limits the force channel value to a certain value. Within an integer range to conform to the MIDI communication protocol standard. Timing Limiter: Limits the timing offset channel value to a preset range. Within the millisecond range, prevent the generation of random rhythms that exceed the limits of human performance.

[0177] : Represents the inverse decoding matrix. This matrix consists of learnable linear projection parameters used to project a high-dimensional acoustic feature space (e.g., 512-dimensional) back into a low-dimensional control parameter space (2-dimensional).

[0178] : Represents the target acoustic feature region obtained through attention mechanism aggregation, i.e., the abstract acoustic representation of "ideal performance".

[0179] : Represents the Hadamard Product, which is an element-wise multiplication operation.

[0180] : Represents the hyperbolic tangent activation function. This item... This constitutes a semantic feature filtering gating system. For the user's query intent vector, This is the gating weight matrix. The gating mechanism outputs within a range based on the user's specific intent (e.g., "stronger explosive power"). The mask vector between them is dynamically enhanced. Select the feature dimensions that are relevant to the intent (such as the energy dimension) and suppress irrelevant dimensions; It is a learnable bias vector, and its role is to provide a baseline adjustment capability for the gating mechanism.

[0181] : Represents the instrument's inherent bias vector. This parameter is used to compensate for the physical differences in touch response between different piano mechanical structures (such as grand pianos and upright pianos). In practical applications, this vector can be initialized by collecting calibration data from a specific piano to ensure that the generated control parameters can reproduce the expected auditory effect on the target instrument.

[0182] In summary, by introducing a semantic gating mechanism and a physical constraint operator, this formula achieves precise reverse decoupling from "abstract auditory perception" to "concrete tactile perception," ensuring that the generated performance suggestions not only meet the needs of emotional expression but also have practical operability.

[0183] Finally, quantized performance control parameters are generated. The quantized performance control parameters are specifically represented as a time-series matrix containing two independent channels, where the first channel is the touch velocity value conforming to the Musical Instrument Digital Interface (MIDI) standard. The second channel is the note timing offset (TimingOffset, ms), measured in milliseconds. Through this step, the system completes the final transformation from "implicit emotional intent" to "explicit technical indicators."

[0184] Example 6, as Figure 6 As shown, the piano music expression teaching system based on affective computing specifically includes: a music semantic association network construction module, a cross-modal feature encoding and generation module, a cross-modal joint semantic space construction module, and an affective expression retrieval and inverse parameter parsing module, wherein;

[0185] The music semantic association network construction module is used to acquire multimodal music data containing musicological text data and demonstration audio data, perform entity extraction and relation parsing on the musicological text data, establish logical connections between emotional description entities and performance technique entities, and construct a music semantic association network.

[0186] The cross-modal feature encoding and generation module is used to encode the music semantic association network as constrained prior information, generate a structured semantic feature vector containing topological information, and perform acoustic signal analysis on the demonstration audio data to generate an acoustic physical feature vector that matches the dimension of the structured semantic feature vector.

[0187] The cross-modal joint semantic space construction module is used to calculate the semantic similarity between structured semantic feature vectors and acoustic physical feature vectors. By minimizing the distance between semantically corresponding samples and maximizing the distance between semantically non-corresponding samples, a cross-modal joint semantic space is established to achieve logical alignment between text semantics and acoustic features within this cross-modal joint semantic space.

[0188] The emotion expression retrieval and inverse parameter parsing module is used to respond to the emotion expression query command for the target song, retrieve the target acoustic feature region that matches the vector representation of the emotion expression query command in the cross-modal joint semantic space, perform inverse parameter parsing on the target acoustic feature region, and output quantized performance control parameters, which include the touch force value and note duration offset.

[0189] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A piano music expression teaching method based on affective computing, characterized in that, Includes the following steps: Step S1: Obtain multimodal music data containing musicological text data and demonstration audio data; perform entity extraction and relation parsing on the musicological text data; establish logical connections between emotion description entities and performance technique entities; and construct a music semantic association network. Step S2: Using the music semantic association network as constraint prior information, feature encoding is performed on the musicological text data to generate a structured semantic feature vector containing topological structure information. At the same time, acoustic signal analysis is performed on the demonstration audio data to generate an acoustic physical feature vector that matches the dimension of the structured semantic feature vector. Step S3: Calculate the semantic similarity between the structured semantic feature vector and the acoustic physical feature vector. By minimizing the distance between semantically corresponding samples and maximizing the distance between semantically non-corresponding samples, a cross-modal joint semantic space is established to achieve logical alignment between text semantics and acoustic features in this cross-modal joint semantic space. Step S4: In response to the emotion expression query command for the target piece, retrieve the target acoustic feature region that matches the vector representation of the emotion expression query command in the cross-modal joint semantic space, and perform inverse parameter parsing on the target acoustic feature region to output quantized performance control parameters, which include the touch force value and the note duration offset.

2. The piano music expression teaching method based on affective computing according to claim 1, characterized in that, In step S1, the specific steps of performing entity extraction and relation parsing on the musicological text data include: First, a pre-set high-dimensional music ontology vector basis space is loaded, and text fragments in the musicology text data are mapped to the high-dimensional music ontology vector basis space. The musicology information density score is determined by calculating the projection distance between the text semantic vector and the music ontology basis vector. The musicology information density score is compared with a preset density threshold, and text fragments with scores lower than the density threshold are removed to obtain high-density text data. Subsequently, a hierarchical multi-task extraction architecture is run for high-density text data. Entity categories are pre-divided into macro-sentiment level, meso-structure level, and micro-technique level. During the extraction process, a hierarchical dependency mask constraint mechanism is introduced to calculate the conditional activation probability of an entity belonging to a specific category. This forces the recognition of micro-technique level entities to depend on the contextual information of macro-sentiment level entities, thereby outputting a set of candidate entity nodes with hierarchical attribute labels.

3. The piano music expression teaching method based on affective computing according to claim 2, characterized in that, In step S1, the steps of establishing a logical connection between the emotion description entity and the performance technique entity, and constructing a music semantic association network, specifically include: The entity pairs in the candidate entity node set are traversed. Natural language dependency parsing is used to identify the modification and modified relationships between entity words. At the same time, a music theory logic verification operator is introduced to query the preset music theory rule base to verify the acoustic rationality of the relationship between emotion description entities and performance technique entities. The confidence of logical relationship edges is calculated by fusing syntactic dependency scores and music theory logic consistency scores. Only logical relationship edges that pass the verification are retained. Subsequently, based on the retained logical relationship edges, a global topological energy minimization pruning operation is performed to construct a global semantic consistency potential energy function that reflects the overall stability of the network. The semantic centrality index of each entity node is calculated, and weak connection subgraphs that conflict with the global sentiment theme are automatically removed. The remaining entity nodes and logical relationship edges after pruning are assembled into a structurally stable music semantic association network.

4. The piano music expression teaching method based on affective computing according to claim 3, characterized in that, In step S2, the specific steps of using the music semantic association network as constraint prior information to perform feature encoding on the musicological text data and generate a structured semantic feature vector containing topological information include: A graph-guided semantic encoder is constructed to map words in musicological text data to initial embedding vectors and align them to the corresponding entity nodes in the music semantic association network. Then, a topological semantic diffusion aggregation operation is performed to extract the confidence of logical relationships corresponding to logical connections in the music semantic association network as transit weights. The semantic information of the current entity node in the network neighborhood is aggregated using a graph convolution mechanism. In the topological semantic diffusion aggregation operation, the semantic features of the nodes themselves are extracted using the self-loop feature transformation matrix, and the topological features diffused from the neighboring nodes are extracted using the neighborhood feature transformation matrix. The topological injection strength coefficient is introduced to numerically adjust the degree of correction of the network structure information. The fused features are processed by a nonlinear activation function to generate a structured semantic feature vector containing rigorous topological logic.

5. The piano music expression teaching method based on affective computing according to claim 4, characterized in that, In step S2, the specific steps of performing acoustic signal analysis on the demonstration audio data and generating acoustic physical feature vectors that are adapted to the dimensions of the structured semantic feature vectors include: A cross-modal technique prior distillation mechanism is adopted to retrieve semantic embedding vectors belonging to performance technique entities from the music semantic association network, and calculate the aggregate mean of semantic embedding vectors to generate global technique prior vectors. Subsequently, the global technique prior vector is used as a gating signal to constrain the acoustic analysis process of the demonstration audio data. Specifically, the dot product correlation operation is performed between the Mel spectrogram features of the demonstration audio data and the global technique prior vector after cross-modal projection to generate a heatmap of attention indicating key frequency band regions. The heatmap of attention is then used to weight and mask the original Mel spectrogram features, forcing the convolutional neural network to focus on the time-frequency region that is strongly related to the performance technique. Finally, the fully connected layer maps and outputs an acoustic physical feature vector whose dimension is strictly consistent with the structured semantic feature vector.

6. The piano music expression teaching method based on affective computing according to claim 5, characterized in that, In step S3, the step of calculating the semantic similarity between the structured semantic feature vector and the acoustic physical feature vector specifically includes: First, a joint manifold space projection operation is performed, introducing the text manifold projection matrix and the acoustic manifold projection matrix as nonlinear mapping heads, which respectively map the structured semantic feature vector and the acoustic physical feature vector onto the shared unit hypersphere. Subsequently, a semantic acoustic consistency metric function is defined and computed. During the computation, a bilinear interaction tensor is introduced to capture the second-order interaction relationship between each dimension of the feature vector. The semantic similarity score, which indicates the degree of sample matching, is obtained by calculating the bilinear product between the projected structured semantic feature vector and the projected acoustic physical feature vector, and dividing the bilinear product by the modulus product adjusted by the temperature scaling factor.

7. The piano music expression teaching method based on affective computing according to claim 6, characterized in that, In step S3, the steps of establishing a cross-modal joint semantic space by minimizing the distance between semantically corresponding samples and maximizing the distance between semantically non-corresponding samples specifically include: Perform spatial alignment operations and construct a topology-adaptive contrastive loss function. Utilize the music semantic association network to calculate the graph geodesic distance between any two entity nodes and transform the graph geodesic distance into semantic neighborhood weights. In the process of calculating the topological adaptive contrastive loss function, the semantic similarity score of semantically corresponding positive sample pairs is maximized, and for semantically non-corresponding negative sample pairs, an exponential function containing a semantic decay feature constant is constructed to construct a topological repulsion factor. The repulsion strength of negative sample pairs in the vector space is dynamically adjusted according to the graph geodesic distance, so that negative sample pairs that are closer in distance in the music semantic association network are subjected to less repulsion. Finally, by minimizing the total topological adaptive contrastive loss, the vector distribution in the joint space is forced to conform to the topological logic of the music semantic association network, thereby establishing a cross-modal joint semantic space.

8. The piano music expression teaching method based on emotion computing according to claim 7, characterized in that, In step S4, the step of retrieving the target acoustic feature region in the cross-modal joint semantic space that matches the vector representation of the emotion expression query instruction in response to the emotion expression query instruction for the target track specifically includes: A graph-guided semantic encoder is reused to convert the received emotional expression query instructions for the target track into a high-dimensional query intent vector. Subsequently, a semantic resonance region scanning operation is performed in the cross-modal joint semantic space to calculate the semantic resonance degree between the query intent vector and each pre-stored acoustic physical feature vector in the space. The multi-head attention mechanism is used to weight and aggregate the acoustic physical feature vectors that present high resonance degree, thereby locking in the target acoustic feature region that best represents the emotional expression query instruction.

9. The piano music expression teaching method based on affective computing according to claim 8, characterized in that, In step S4, the specific steps of performing inverse parameter analysis on the target acoustic feature region and outputting quantized performance control parameters include: An adaptive gating inverse decoupling network is constructed, and the query intent vector is used as the driving signal input to the gating weight matrix. After processing by the hyperbolic tangent activation function, a semantic feature filtering gating for a specific performance dimension is generated. Using semantic feature filtering gates, element-wise multiplication operations are performed on the target acoustic feature region to extract latent feature components that are strongly related to performance control from the abstract acoustic features. The latent feature components are then input into the inverse decoding matrix. The high-dimensional feature space is mapped back to the low-dimensional control parameter space through matrix projection transformation, and the instrument's inherent deviation vector is superimposed. Subsequently, a physical constraint transformation operator is introduced to post-process the mapped values, forcing the dimensionless neural network output values ​​to be mapped into the physically feasible domain of piano performance biomechanics, and finally generating quantized performance control parameters. These quantized performance control parameters are specifically represented as a time series matrix containing two independent channels, where the first channel is the touch force value conforming to the instrument digital interface standard, and the second channel is the note value offset in milliseconds.

10. The piano music expression teaching system based on affective computing is applied to the piano music expression teaching method based on affective computing as described in any one of claims 1-9, characterized in that, Specifically, it includes: a music semantic association network construction module, a cross-modal feature encoding and generation module, a cross-modal joint semantic space construction module, and an emotion expression retrieval and inverse parameter parsing module, among which; The music semantic association network construction module is used to acquire multimodal music data containing musicological text data and demonstration audio data, perform entity extraction and relation parsing on the musicological text data, establish logical connections between emotional description entities and performance technique entities, and construct a music semantic association network. The cross-modal feature encoding and generation module is used to encode the music semantic association network as constrained prior information, generate a structured semantic feature vector containing topological information, and perform acoustic signal analysis on the demonstration audio data to generate an acoustic physical feature vector that matches the dimension of the structured semantic feature vector. The cross-modal joint semantic space construction module is used to calculate the semantic similarity between structured semantic feature vectors and acoustic physical feature vectors. By minimizing the distance between semantically corresponding samples and maximizing the distance between semantically non-corresponding samples, a cross-modal joint semantic space is established to achieve logical alignment between text semantics and acoustic features within this cross-modal joint semantic space. The emotion expression retrieval and inverse parameter parsing module is used to respond to the emotion expression query command for the target song, retrieve the target acoustic feature region that matches the vector representation of the emotion expression query command in the cross-modal joint semantic space, perform inverse parameter parsing on the target acoustic feature region, and output quantized performance control parameters, which include the touch force value and note duration offset.