Method and system for intelligent parameterization extraction and representation of characteristics of opera singing tunes

By preprocessing and structural segmenting the audio data of opera singing, parameters of the vocal structure layer, vocal embellishment layer, and dynamics layer are extracted, and a unified parameter model is constructed. This solves the problem of distinguishing between opera genres and schools in the analysis of opera singing features, realizes feature extraction and representation across opera genres and schools, and improves the accuracy and applicability of the analysis.

CN122369504APending Publication Date: 2026-07-10LANZHOU YINQIAO CULTURAL COMMUNICATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LANZHOU YINQIAO CULTURAL COMMUNICATION CO LTD
Filing Date
2026-04-27
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing techniques for analyzing the characteristics of traditional Chinese opera singing styles cannot distinguish between different opera genres and schools, resulting in inconsistent feature extraction effects. They are unable to conduct accurate analysis at the level of vocal style, vocal embellishment, and emotional expression, making it difficult to meet the research and application needs across opera genres and schools.

Method used

By acquiring audio data of opera singing, preprocessing and structural segmentation are performed, and parameter vectors of the cavity structure layer, embellishment layer and dynamics layer are extracted. A unified parameter model is constructed, and the feature output is adaptively adjusted according to downstream tasks to achieve feature extraction and representation across opera genres and schools.

Benefits of technology

It enables dynamic parameter adjustment of opera singing characteristics, improves applicability and analysis accuracy across different opera genres and schools, and supports various application scenarios such as opera musicology research, school identification, and emotion computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369504A_ABST
    Figure CN122369504A_ABST
Patent Text Reader

Abstract

This invention provides an intelligent parameterized extraction and representation method and system for opera singing features, relating to the field of speech analysis and recognition technology. By constructing a configuration mapping table and rule engine containing genre labels, school labels, and downstream task labels, a configuration vector containing parameter set switches, weight values, and output granularity information is automatically generated before analysis begins. Based on downstream labels such as classification, comparison, teaching, and AI generation, the weight values ​​and output granularity of the set parameters are automatically adjusted. Using the configuration vector as a constraint, corresponding feature extraction is performed, decomposing opera singing features into three levels for extraction and fusion. Parameter vectors for the vocal structure layer and embellishment layer are extracted, and dynamic layer parameters are extracted by calculating emotional tension curves. A structured fusion is used to form a unified parameter model. The system can select to output a complete model or a subset based on the downstream task, achieving a targeted parameter generation strategy for different downstream tasks, improving accuracy and adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech analysis and recognition technology, specifically to a method and system for intelligent parameterized extraction and representation of opera singing features. Background Technology

[0002] As a core expressive element of traditional Chinese opera, opera singing possesses highly complex musical structures and performance techniques. These include pitch gliding and embellishment within a non-fixed pitch system, strong rhythmic flexibility and free meter, melodic deformation based on linguistic tone and emotional expression, and stable yet difficult-to-quantify stylistic characteristics formed by different opera genres, schools, and roles. Currently, the digital analysis of opera singing primarily relies on manual annotation and musicological analysis methods, or employs general audio feature extraction techniques oriented towards Western music systems. Existing technologies suffer from the following technical problems: Problem 1: In existing opera singing feature analysis techniques, the same feature extraction method is applied indiscriminately to opera genres with vastly different acoustic characteristics, such as Peking Opera, Qinqiang Opera, and Huangmei Opera. This results in inconsistent feature extraction results, and the different styles of different schools cannot be reflected in feature extraction. Consequently, there is insufficient research and analysis of cross-genre operas, and schools cannot be clearly and automatically distinguished. Furthermore, since different tasks such as classification, teaching, and AI generation have different requirements for feature granularity and output format, the existing methods use a fixed output format, which is difficult to adapt, thus limiting the accuracy and practicality of the analysis results. Problem 2: Existing technologies cannot incorporate the stylized common features, personalized vocal techniques, and emotional tension changes of traditional Chinese opera into a unified parametric representation system. Simply extracting and classifying opera features results in the mixed extraction of features at different levels, and the inability to decouple and analyze them leads to the inability to accurately describe and judge the vocal style, vocal techniques, and emotional aspects of the singing in opera research. This makes it difficult to simultaneously serve multiple application scenarios such as opera musicology research, genre identification, and emotion computing. Summary of the Invention

[0003] To achieve the above objectives, the present invention provides the following technical solution: a method for intelligent parameterized extraction and representation of opera singing features, the method comprising: The original opera singing audio is acquired and preprocessed to obtain the singing audio data; Obtain the input genre tags, style tags, and downstream task tags, and generate a configuration vector by querying a pre-stored configuration mapping table; Structural segmentation is performed on the vocal audio data to detect breathing points, cavity boundaries, and melodic boundaries, resulting in a time-stamped sequence containing musical phrases, cavities, and melodic patterns. Based on the configuration vector, the cavity grid layer parameter vector and the embellishment layer parameter vector are extracted from the audio segments corresponding to the time stamp sequence in the vocal audio data; Based on the configuration vector, vocal audio data, time stamp sequence, and cavity layer parameter vector, the pitch tension sequence, dynamic tension sequence, and rhythmic elasticity tension sequence are calculated and synthesized into an emotional tension curve. Dynamic statistical features are extracted from the emotional tension curve to obtain the dynamic parameter vector. The cavity layer parameter vector, the cavity lubrication layer parameter vector, the dynamic parameter vector, and the emotional tension curve are structurally fused to obtain a unified parameter model; Based on the downstream task labels, the corresponding output feature vectors are selected from the unified parameter model.

[0004] On the other hand, this application also proposes an intelligent parameterized extraction and representation system for opera singing features, used to implement the aforementioned intelligent parameterized extraction and representation method for opera singing features. The system includes: The audio acquisition module is used to acquire the original opera singing audio and preprocess it to obtain singing audio data; The parameter configuration module is used to obtain the input genre tags, style tags, and downstream task tags, and generate configuration vectors by querying the pre-stored configuration mapping table; The vocal structure segmentation module is used to perform structural segmentation on vocal audio data, detect breathing points, cavity boundaries and melodic boundaries, and obtain a time-stamped sequence containing musical phrases, cavities and melodic patterns. The feature extraction module is used to extract the cavity grid layer parameter vector and the embellishment layer parameter vector from the audio segments corresponding to the time stamp sequence in the vocal audio data based on the configuration vector; The dynamics parameter module is used to generate an emotional tension curve based on the configuration vector, vocal audio data, time stamp sequence, and cavity grid parameter vector, and to extract dynamic statistical features from the emotional tension curve to obtain the dynamics parameter vector. The parameter fusion module is used to structurally fuse the cavity layer parameter vector, the lubrication layer parameter vector, the dynamic parameter vector, and the emotional tension curve to obtain a unified parameter model; The adaptive output module is used to select and generate corresponding output feature vectors from the unified parameter model based on the downstream task labels.

[0005] This invention provides a method and system for intelligent parameterized extraction and representation of opera singing features. It has the following beneficial effects: 1. This invention constructs a configuration mapping table and rule engine containing genre tags, style tags, and downstream task tags. Before analysis begins, it automatically generates a configuration vector containing parameter set switches, weight values, and output granularity information. Based on the downstream tags (classification, comparison, teaching, AI generation, etc.), it automatically adjusts the weight values ​​and output granularity of the set parameters to achieve adaptive information configuration. Using the configuration vector as a constraint, it performs corresponding feature extraction, ensuring that the extracted features correspond one-to-one with the needs of the downstream tags. This solves the inherent problems of one-size-fits-all and mixed feature extraction in existing technologies. It allows the same method to adapt to different application scenarios without modifying the algorithm. It can dynamically adjust parameters and extract features for different genres, styles, and tasks, effectively improving the accuracy of parameter configuration. This is superior to comparison schemes using fixed parameter sets, achieving applicability and flexibility across genres, styles, and multiple tasks, and increasing the accuracy of feature extraction and analysis.

[0006] 2. This invention employs a three-level extraction and fusion approach to decompose and represent the characteristics of traditional Chinese opera singing. By extracting parameter vectors from the vocal structure layer, it reflects the stylized common structures of different opera genres and schools, exhibiting high stability. By extracting parameter vectors from the embellishment layer, it reflects the individualized singing techniques of actors, demonstrating high variability. Finally, by calculating emotional tension curves, it extracts parameters from the dynamics layer, quantifying abstract emotional fluctuations into a computable time series. These three layers of parameters are structurally fused to form a unified parameter model. The model can be output as a complete model or a subset depending on the downstream task. This allows traditional Chinese opera musicology research to focus on vocal structure layer parameters, genre identification to combine vocal structure and embellishment layers, and emotion calculation and AI generation to rely on dynamics layer parameters. This represents a leap from single-feature to multi-dimensional decoupled fusion, enabling targeted parameter generation strategies for different downstream tasks. It improves the accuracy of opera genre classification, the precision of teaching feedback, and the depth of understanding of emotion analysis, achieving significant progress in feature representation completeness and application versatility. Attached Figure Description

[0007] Fig. 1 This is a flowchart illustrating the steps of the intelligent parameterized extraction and characterization method for opera singing features of the present invention. Fig. 2 This is a data transmission flowchart of the intelligent parameterized extraction and characterization method for opera singing features of the present invention; Fig. 3 This is an architecture diagram of the intelligent parameterized extraction and characterization system for opera singing features of the present invention. Detailed Implementation

[0008] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0009] like Figs. 1-2 As shown, an intelligent parameterized extraction and representation method for opera singing features is presented, comprising the following steps: Step S100: Acquire the original opera singing audio and preprocess it to obtain singing audio data. The singing audio data reflects the pure singing signal form of the original opera singing audio after a series of standardization and purification processes. It removes irregular silent segments, weak noise interference, and channel inconsistencies caused by environment, equipment, or singing intervals in the original recording, and avoids feature deviations caused by underlying signal quality problems.

[0010] Step S200: Obtain the input genre label, style label, and downstream task label. Generate a configuration vector by querying a pre-stored configuration mapping table. The genre label guides the system to enable specific analysis rules strongly related to the genre. The style label dynamically adjusts the sensitivity or weight of the embellishment parameter set switch in subsequent extraction. The downstream task label controls the granularity and detail of the final output information. The configuration vector is a set of specific, quantified machine instructions compiled from the semantic information of the above three labels, including which parameter sets to enable, the weight values ​​of each parameter set, and the level of detail in the output.

[0011] Step S300 involves structural segmentation of the vocal audio data, detecting breath points, phrasal boundaries, and melodic boundaries to obtain a time-stamped sequence containing musical phrases, phrasal sections, and melodic patterns. This structural segmentation of the vocal audio data simulates the process in music theory of dividing a continuous performance into musical phrases, then into phrasal sections, and finally into melodic patterns. The resulting time-stamped sequence contains the start and end times of each musical phrase, each phrasal section, and each melodic pattern, ensuring that the phrasal layer parameters, embellishment layer parameters, and dynamic parameters can be accurately calculated from the correct audio segments, achieving refined and aligned analysis of the performance content.

[0012] Step S400: Based on the configuration vector, extract the vocal style parameter vector and embellishment parameter vector from the audio segments corresponding to the time-stamped sequence in the vocal audio data. The extracted vocal style parameter vector is used to achieve automatic classification and clustering analysis of length at the opera genre and school level. The extracted embellishment parameter vector represents the personalized and improvisational embellishments made by actors on top of the stylized main body, and is used to quantify the most significant stylistic differences between different singers of the same aria, providing specific numerical indicators for technique evaluation in opera teaching or fine comparison within a school.

[0013] Step S500: Based on the configuration vector, vocal audio data, time-stamped sequence, and melody layer parameter vector, calculate the pitch tension sequence, dynamic tension sequence, and rhythmic elasticity tension sequence, and synthesize them into an emotional tension curve. Extract dynamic statistical features from the emotional tension curve to obtain a dynamic parameter vector. The emotional tension curve reflects the dynamic process of emotional tension and release that continuously changes over time in opera singing, perceptible to the audience. It transforms the physical aspects of pitch deviation, dynamic changes, and rhythmic elasticity into a single psychoacoustic proxy curve through weighted fusion, visualizing and temporally sequenced abstract emotional fluctuations. It can intuitively show the complete emotional narrative arc of a performance, from calm to fervor, and then back to calm. The dynamic parameter vector is a macroscopic statistical feature extracted from the emotional tension curve, condensing the emotional expression of the entire performance into a set of numerical labels that can be compared and classified, facilitating qualitative or quantitative analysis of the emotional type of the performance.

[0014] Step S600 involves structurally fusing the cavity layer parameter vector, embellishment layer parameter vector, dynamic parameter vector, and emotional tension curve to obtain a unified parameter model. This unified parameter model reflects a comprehensive, multi-layered, and structured digital twin of a traditional Chinese opera aria, encompassing everything from micro-techniques to macro-structure and emotional dynamics. Depending on the needs of downstream tasks, it can selectively include deviation vectors from the standard template and complete emotional curves. Serving as a core hub connecting underlying audio analysis with upper-level diverse applications, it provides a complete yet decoupled data set. Whether it's an algorithm requiring abstract statistical features for classification, a model needing frame-by-frame temporal information for AI synthesis, or an application needing to compare deviations for teaching, the most suitable data subset can be extracted from the unified model as needed.

[0015] Step S700: Based on the downstream task labels, select and generate corresponding output feature vectors from the unified parameter model. By selectively generating output feature vectors, the massive internal unified parameter model is intelligently trimmed and transformed according to the initially input task instructions, achieving a mapping from complete information to precise information, avoiding redundancy or omissions caused by a one-size-fits-all output.

[0016] In this embodiment, the steps of acquiring the original opera singing audio and preprocessing it to obtain the singing audio data include: Step S101: The acquired original opera singing audio is converted into a mono signal with a uniform sampling rate of 22.05kHz. The sampling rate of 22.05kHz is sufficient to capture the fundamental frequency and main overtones of the opera singing within the range of human hearing, and it is a commonly used standard in the field of signal processing. When performing mono signal conversion, the original stereo or multi-channel signal of the original opera singing audio is first read, and then the left and right channel data are added together and divided by two to form a mono sequence. The sampling point density of the mono sequence is adjusted from the original sampling rate to a uniform interval of 22.05 kHz through a digital resampling filter to eliminate phase interference between channels and unify the time resolution of the data.

[0017] The energy of segments within a preset time window of 10ms is calculated. The signal in the mono sequence is framed and windowed (Hamming window) to obtain a frame sequence. A fast Fourier transform is performed on each frame to calculate its power spectral density. Based on the power spectrum, the sub-band energy is calculated in the frequency band of 50Hz-2000Hz. Segments with energy below the first energy ratio of the maximum energy and a duration exceeding the first time threshold are cut off. Long silences and breathing intervals before and after the performance are removed as ineffective performance areas to avoid interference from silent segments with subsequent breathing point detection and phrase segmentation. The amplitude distribution of the cut segments is statistically analyzed, and the amplitude threshold is set to the first multiple of the maximum amplitude. Sampling points in the audio below the amplitude threshold are set to zero to further refine the suppression of weak background noise within the effective singing segments and improve the purity of the signal.

[0018] The original opera singing audio is an unprocessed initial audio waveform file obtained directly from live recordings, historical recording tapes, digital recording files, etc., and is generally uploaded by the user, recorded in real time by a microphone array, or read from a database. In a standard configuration, the first energy ratio is generally set between 1% and 5%, and in this embodiment, the first energy ratio is 1%. The first time threshold is set between 0.3 seconds and 0.8 seconds, and in this example, the first time threshold is 0.5 seconds. The first amplitude multiplier is set between 1.2 times and 2.0 times, and in this example, the first amplitude multiplier is 1.5 times. The specific value can be fine-tuned according to the background noise level of the recording environment.

[0019] Step S102: Adjust the maximum amplitude of the entire audio to 0.95 times full scale to obtain the vocal audio data. Based on the general principle of preventing clipping distortion in digital audio processing, adjusting the maximum amplitude of the entire audio to 0.95 times full scale, i.e., retaining 5% headspace, effectively avoids instantaneous overshoot introduced by subsequent filtering or resampling operations. During adjustment, first scan the sequence of the entire vocal audio data and find the amplitude value with the largest absolute value among all sampling points as the peak value; then calculate a global gain coefficient, which is equal to the target amplitude (0.95 multiplied by the maximum possible encoded value) divided by the peak value; finally, multiply each sampling point in the audio data by this global gain coefficient to achieve linear scaling of the overall amplitude, ensuring that vocal segments recorded at different loudness levels have a consistent average loudness benchmark before entering the analysis module. In step S102, the global digital gain of the vocal audio data purified in step S101 is adjusted so that the largest peak in the entire audio waveform reaches exactly 95% of the full scale of the digital system. This fully utilizes the dynamic range of the digital signal and improves the signal quantization accuracy while retaining a small safety margin to prevent unexpected peak clipping distortion during processing.

[0020] The final vocal audio data includes a mono digital signal sequence with a sampling rate of 22.05 kHz. Low-energy silence segments in the sequence have been removed, sampling points below the amplitude threshold have been set to zero, and the maximum absolute amplitude of the entire sequence has been normalized to 95% of full scale.

[0021] In this embodiment, the steps of obtaining the input genre tags, style tags, and downstream task tags, and generating a configuration vector by querying a pre-stored configuration mapping table include: Step S201 uses the obtained combination of opera genre tags, style tags, and downstream task tags as an index to search the pre-stored configuration mapping table. The configuration mapping table is a structured data table pre-stored in the system. This table includes columns for opera genre tags, style tags, downstream task tags, and corresponding configuration vector columns. The configuration vector columns store the pre-calculated parameter set switch states, weight values, and output granularity for a specific combination. The configuration mapping table is based on a summary of extensive opera acoustic analysis and music theory knowledge; that is, it identifies which features are most critical for a particular opera genre or style when completing a certain type of task, thus forming the corresponding configuration vector columns. Opera genre tags are used to identify the art form to which the opera singing style belongs, including Qinqiang, Peking Opera, Henan Opera, Yue Opera, etc. These tags are usually provided by the user when initiating an analysis request through a graphical interface drop-down menu, command-line parameters, or direct input via API, and are used to guide the activation of feature analysis logic strongly related to a specific opera genre. The genre label identifies the singer's style or role, including female roles, male roles, painted-face roles, Mei School, Cheng School, etc. It is provided by the user upon request and can be empty. It is used to adjust parameter weights based on genre characteristics when extracting the embellishment layer parameter vector. The downstream task label identifies the intended use of the analysis results, including classification, comparison, teaching, AI generation, etc. It is input by the user and controls the content structure, data granularity, and whether bias feedback is included in the final output feature vector.

[0022] If an exact match exists, the corresponding configuration vector is read directly; if no exact match exists, the closest configuration is matched in the order of priority: genre first, style second, and downstream task last, and a configuration vector is generated according to preset configuration rules: the preset configuration rules include the following rules: All opera genres have their basic vocal parameters (range, mode, and rhythm) enabled; Qinqiang opera additionally has its characteristic parameters (bitter and joyful tones) enabled. Since the stylized commonalities of all opera genres are reflected in their vocal range, mode organization, and rhythmic patterns, which are fundamental dimensions for describing any aria, this ensures that at least one set of basic parameters describing the vocal structure can be extracted under any circumstances, providing a minimum amount of common information for subsequent analysis. By enabling the bitter and joyful tones, Qinqiang opera is accurately characterized, ensuring that the analysis of Qinqiang opera retains its genre-specific characteristics.

[0023] If the style label is not empty, the weight of the embellishment parameter will be increased by 0.2. If the style label is "Dan" (female role), the weight of the vibrato parameter will be increased by an additional 0.1. Style is mainly reflected in personalized embellishment techniques. When the user specifies a style, it means that the need for analyzing style differences is enhanced. By dynamically amplifying the influence of embellishment features in subsequent integration, the results can better reflect the characteristics of the style. Furthermore, Dan singing usually uses vibrato more frequently and more obviously as a style marker than Sheng or Jing singing. Further refinement of weight adjustment is needed to adapt to the performance characteristics of Dan.

[0024] If the downstream task is labeled "teaching," then the output granularity is set to the pitch pattern level, and the required deviation value is marked. Teaching scenarios require the finest granularity to pinpoint specific positional deviations in students' singing, providing detailed feedback data on pitch patterns.

[0025] If the downstream task is labeled as classification, the output granularity is set to sentence level and the statistical features need to be compressed. Classification tasks rely more on macroscopic statistical information than detailed temporal information, and feature compression helps improve classifier efficiency and generalization ability, resulting in a compact feature vector suitable for machine learning input.

[0026] The preset configuration rules are a set of logical judgment statements used by the system to automatically and dynamically generate a reasonable configuration vector when there is no record in the configuration mapping table that completely matches the input label. The preset configuration rules define the priority order of the genre, style, and downstream task, as well as the specific parameter adjustment strategy. Based on the knowledge priority settings in the field of opera music, the genre determines the most basic acoustic paradigm and therefore has the highest priority; the style provides style fine-tuning within the genre framework and has the next highest priority; the downstream task determines the final output organization form and has the next lowest priority, ensuring that even when faced with undefined label combinations.

[0027] Step S202: The configuration vector includes a parameter set switch list, a weight value list, and output granularity. The parameter set switch list includes a cavity parameter set switch, an embellishment parameter set switch, and a dynamics parameter set switch, with each parameter set switch assigned a weight value. The output granularity includes sentence level, cavity level, and tone pattern level. The parameter set switch list is used to control whether the three major parameter categories (cavity, embellishment, and dynamics) in the feature extraction module are enabled. When the cavity parameter set switch is enabled, the system will execute steps S401 to S402 to extract stable features such as vocal range, mode, and tempo. When the embellishment parameter set switch is enabled, the system will execute steps S403 to S404 to extract decorative features such as glissando, vibrato, appoggiatura, and euphony. When the dynamics parameter set switch is enabled, the system will execute step S500 to calculate the emotional tension curve and dynamics parameter vector. The weight value is used in step S502 to generate the emotional tension curve. It is used to sum the pitch tension, dynamic tension and rhythm tension in a weighted manner to reflect the differences in importance of different tension components in different tasks or genres. Each parameter set switch is assigned a weight value.

[0028] Output granularity indicates the time unit level to which the final output parameters are attached, determining whether features are organized by phrase, melody, or pattern. Output granularity includes three levels: phrase-level, melody-level, and pattern-level, used to control the fineness of the analysis results. Coarse granularity is suitable for rapid classification, while fine granularity is suitable for teaching feedback. The output granularity is selected and adjusted based on the detail requirements of the downstream task: classification tasks typically only require phrase-level statistical features, while teaching tasks require pattern-level point-by-point comparisons. Specifically, phrase-level refers to using a complete musical phrase as the basic unit for parameter calculation and output, providing a macro-level overview of the entire singing passage. Melody-level refers to using a melody segment (a smaller unit within a phrase separated by breaths or pauses) as the basic unit, used for analyzing melodic direction at a medium granularity. Pattern-level uses a pattern (an alternating unit of fundamental frequency stability and variation within a melody segment) as the basic unit, used to capture the most subtle changes in singing technique. The output granularity determines whether to detect and record pattern boundaries; pattern boundary detection is only necessary when the output granularity is pattern-level.

[0029] In this embodiment, the implementation steps of structural segmentation of the vocal audio data, detection of breath points, phrasal boundaries, and melodic boundaries, to obtain a time-stamped sequence containing musical phrases, phrasal sections, and melodic patterns include: Step S301: Divide the vocal audio data into a frame sequence with a specified frame length of 25ms and a specified frame shift of 10ms. For each frame, extract all sampling points within that frame. Since the instantaneous power of the sound is proportional to the sum of the squares of the sampling points, square the value of each sampling point. Then calculate the sum of the squares of the sampling points within each frame to obtain the short-time energy of that frame. Connect the calculation results of all frames in chronological order to form an initial short-time energy sequence. Perform a moving average smoothing process on the initial short-time energy sequence to obtain the final short-time energy envelope curve. The short-time energy envelope curve is a smooth curve that varies with time. The vertical axis represents the average intensity of the audio signal within a local time window and consists of a series of continuous short-time energy values, each corresponding to an analysis frame. The short-time energy envelope curve reflects the macroscopic contour of the volume fluctuations over time during the performance; strong areas correspond to accented or exciting parts, while weak areas correspond to gentle or breathing parts.

[0030] The specified frame length is the duration of the audio segment analyzed in each step, which is 25 milliseconds in this example. A longer frame length results in higher frequency resolution, allowing for more precise differentiation of two similar pitches, but lower temporal resolution, making it difficult to capture rapid changes. The specified frame shift is the time interval between the start points of two adjacent analysis frames, which is 10 milliseconds in this example. This determines the temporal sampling density of the feature sequence; a smaller frame shift results in higher temporal resolution but also higher computational cost. The specified frame length and frame shift are based on classic empirical settings in speech / music signal processing. A 25-millisecond frame length can cover at least two fundamental frequency cycles (for bass frequencies), and a 10-millisecond frame shift provides sufficient temporal overlap to ensure smooth feature sequences. Generally, the frame length is set between 20-30 milliseconds, and the frame shift between 5-15 milliseconds. In this embodiment, a 25-millisecond frame length and a 10-millisecond frame shift are chosen to balance computational efficiency and analysis accuracy.

[0031] Step S302: In the short-time energy envelope curve, the location where the energy drops sharply by more than 50% of the second energy ratio and lasts for more than 30ms of the second time threshold is marked as a breathing point. The interval between adjacent breathing points is a musical phrase, and the start and end times of each musical phrase are recorded. By detecting breathing points, the moment when the singer breathes or pauses briefly can be identified, thereby segmenting continuous audio into independent musical phrases, which is the basis for analyzing melodic grammar.

[0032] Within each musical phrase, the fundamental frequency change rate between adjacent frames is calculated. Positions where the absolute value of the fundamental frequency change rate exceeds twice the average change rate by a first preset multiple and is accompanied by a decrease in energy are marked as segment boundaries. The start and end times of each segment are recorded. By detecting segment boundaries, a musical phrase is further subdivided into smaller melodic units, typically corresponding to a lyric or a short melodic fragment, which helps in analyzing the rhythm and pitch organization within the phrasing. The fundamental frequency change rate represents the drastic change in the fundamental frequency value between two adjacent frames, usually measured in cents or semitones per millisecond. It quantifies the smoothness or abruptness of the melody, thus helping to locate segment boundaries where the change rate suddenly increases and melodic boundaries where the change rate suddenly decreases. The calculation process for the fundamental frequency change rate between adjacent frames is as follows: First, the fundamental frequency values ​​of the Nth and N+1th frames are obtained in Hertz. These two Hertz values ​​are converted to logarithmic values ​​(such as semitones or cents) using a conversion formula of 1200 multiplied by a base-2 logarithmic ratio to conform to human logarithmic perception. Then calculate the absolute value of the difference between the two converted values, divide this difference by the specified frame shift time of 10 milliseconds, and obtain the change per millisecond, i.e., the fundamental frequency change rate.

[0033] When the output granularity in the configuration vector is at the pitch pattern level, the alternation points between the fundamental frequency stable segment and the fundamental frequency changing segment are detected within each cavity segment. Intervals where the fundamental frequency change rate is continuously less than 1.5 times the average change rate by a second preset multiple and the duration exceeds a third time threshold of 40ms are marked as stable segments. The alternation points are defined as pitch pattern boundaries, and the start and end times of each pitch pattern are recorded. By detecting pitch pattern boundaries, the alternation points between stable and changing pitch segments are identified within the cavity segment, thus decomposing the performance into the most basic pitch and duration building blocks. This is a prerequisite for analyzing microscopic vocal embellishments such as glissando and vibrato.

[0034] The second energy ratio, second time threshold, first preset multiple, second preset multiple, and third time threshold are set based on statistical analysis of a large amount of manually labeled data of opera singing to maximize segmentation accuracy. The general setting range is: second energy ratio 40%-60%, second time threshold 20-50 milliseconds, first preset multiple 1.5-2.5 times, second preset multiple 1.2-1.8 times, and third time threshold 30-60 milliseconds.

[0035] Step S303: Combine the start and end times of each musical phrase, each melody, and each musical pattern into a time-stamped sequence. The time-stamped sequence precisely records the start and end times of each musical phrase, each melody, and each musical pattern in the vocal audio data after structural segmentation. This sequence includes a list of musical phrases, each phrase entry containing the start and end timestamps of that phrase and a list of melody segments belonging to that phrase, and each melody segment entry containing the start and end timestamps of that melody segment and a list of musical patterns belonging to that melody segment.

[0036] In practical applications, the entire calculation process of steps S301, S302, and S303 can be implemented in a general digital signal processing software environment. For example, Python can be used in conjunction with the librosa, pydub, or scipy libraries. Alternatively, MATLAB's audio processing toolbox can be used, or frameworks such as Essentia or JUCE can be used in a C++ environment. In specific implementation, technicians will write code to read the audio array, process each frame in a loop according to the frame length and frame shift described above, calculate the energy and fundamental frequency sequence, and then write logical judgment functions to detect energy drop points, fundamental frequency change rate inflection points, etc., and finally output the time-stamped sequence.

[0037] In this embodiment, the process of extracting the cavity lattice layer parameter vector includes: Step S401: For each musical phrase, extract the fundamental frequency of all frames, find the lowest and highest fundamental frequencies as the lower and upper bounds of the pitch range, and determine the commonly used pitch areas. The fundamental frequency extraction process includes: first, for each frame of audio signal, using an autocorrelation algorithm, calculating the similarity between the frame signal and its delayed version, and finding the delay time corresponding to the maximum similarity. Then, calculating the fundamental frequency based on the delay time, where the fundamental frequency equals the sampling rate divided by the delay time. Finally, using median filtering to smooth the extracted fundamental frequency sequence, removing harmonic or half-frequency errors caused by harmonic interference. Because speech signals have quasi-periodicity, the autocorrelation function will show a peak at the fundamental frequency period, making it a robust fundamental frequency estimation method suitable for vibrato and glissando commonly found in traditional Chinese opera singing. The fundamental frequency is the lowest frequency of vibration of the sound-producing body, measured in Hertz, used to quantify the melodic contour of singing, and is the basis for calculating almost all pitch-related features such as pitch range, mode, interval, glissando, and vibrato.

[0038] When determining the commonly used pitch range, the fundamental frequency values ​​of all frames within the entire singing segment are statistically analyzed. Intervals are divided in semitones, and a fundamental frequency distribution histogram is constructed. From this histogram, the group of continuous intervals with the highest proportion of total frames is identified. The minimum fundamental frequency of this group is used as the lower bound of the commonly used pitch range, and the maximum fundamental frequency is used as the upper bound. In this embodiment, the fundamental frequency range comprising more than 60% of the frames is selected as the commonly used pitch range. By determining the commonly used pitch range, the range of pitches most frequently used by the singer in the singing segment is reflected, rather than their physiological limits. This better represents the stylistic characteristics and difficulty of the singing segment than simply the highest and lowest notes. The horizontal axis of the fundamental frequency distribution histogram represents the fundamental frequency intervals divided by semitone intervals, such as C4, C#4, D4, etc., while the vertical axis represents the number or percentage of analysis frames falling into each interval. Based on the fundamental frequency values ​​extracted from each frame in the entire singing segment, the fundamental frequency distribution histogram reflects the distribution of the singer's dwell time at various pitches, i.e., which pitches are sung more and which are sung less. It includes information such as the pitch concentration trend, the vocal range distribution pattern, and the possible tonic position of the singing segment.

[0039] All fundamental frequencies are mapped to pitch class numbers based on the tonic. A pitch class histogram is generated by statistically analyzing the frequency of each pitch class. The mode type is determined based on the peak patterns of the pitch class histogram. First, the tonic of the aria is determined by finding the highest frequency pitch in the fundamental frequency distribution histogram as a candidate tonic, or by user input, such as the slightly tuned tonic of Qinqiang opera. Then, each fundamental frequency value is converted to a pitch class number relative to that tonic. For example, if the tonic is D4, then D4 itself is mapped to 0, D#4 to 1, C4 to -1, and so on. All fundamental frequencies are replaced with these integer pitch class numbers. Finally, after completing all fundamental frequency mappings, a pitch class histogram is created with the pitch class number relative to the tonic as the x-axis and the percentage of frames each pitch class appears in as the y-axis. Finally, the mode type is determined. All local peaks in the pitch histogram are identified, i.e., pitches higher than both the left and right sides. A threshold is set, for example, a peak height exceeding 1.5 times the average height is considered significant. The number of significant peaks is then counted. If there are 5 significant peaks, roughly corresponding to pentatonic positions, it is determined to be a pentatonic mode; if there are 7 significant peaks, it is determined to be a heptatonic mode. For Qinqiang opera, it is also necessary to additionally detect whether the actual fundamental frequencies of the fa and si notes deviate from the twelve-tone equal temperament (i.e., whether they are too low or too high), outputting a "bitter" or "joyful" mark to complete the mode type determination.

[0040] The fundamental frequency mapping is based on the concept of solfège in music theory, meaning that the same melody played in different keys has the same sequence of pitches relative to the tonic. This allows for direct comparison of singing segments in different keys, eliminating tonal differences and extracting the pure melodic interval structure. The pitch histogram reflects the frequency of each pitch relative to the tonic in a singing segment, clearly revealing the scale structure of the mode, such as a pentatonic scale (with only five significant pitches) or a heptatonic scale. It contains the weight information of each pitch and is the most direct basis for determining the mode type. The peak pattern of the pitch histogram refers to the shape characteristics of its peak distribution. Mode types include pentatonic modes (Gong, Shang, Jiao, Zhi, Yu), heptatonic modes (Qingyue, Yayue, Yanyue), and variations of specific opera genres (such as the bitter and joyful tones of Qinqiang opera). The basis for determining the mode type is the definition of mode in musicology. A mode consists of a specific set of pitches and is used to automatically identify the mode system to which a singing segment belongs, providing key features for opera genre classification.

[0041] Within a musical phrase, the fundamental frequency difference between adjacent melodic patterns is calculated sequentially. Using semitones as units, the frequency of minor seconds, major seconds, and minor thirds is counted, and the distribution of all differences is used to generate an interval distribution histogram. Adjacent melodic patterns refer to two immediately preceding and following melodic patterns within the same melody in a time-stamped sequence. By calculating the fundamental frequency difference between these two patterns, the interval size of the melody is quantified. The statistical distribution of these differences, i.e., the interval distribution histogram, reflects which intervals are frequently used in the singing passage, thus indirectly confirming the modal style and singing techniques. The interval distribution histogram is a statistical graph built with the interval size (in semitones) on the x-axis and the frequency of that interval in the singing passage on the y-axis. It is based on the fundamental frequency difference (absolute value or signed direction) between all adjacent melodic patterns, reflecting the preference for interval leaps in the melody, such as a preference for stepwise motion (minor seconds, major seconds) or leaps (thirds and above). It includes the frequency of various intervals and is an important statistical quantity describing the characteristics of the melodic contour.

[0042] Local peaks in the short-time energy envelope curve are extracted as accent positions. The time interval between adjacent accent positions is calculated to obtain an interval list. Histogram statistics are performed on the interval list, and the interval with the highest frequency among adjacent accent intervals is taken as the basic beat cycle. The rhythmic pattern type is determined based on the pattern of strong accents. The number of basic beat cycles between each strong accent and the next strong accent is calculated. If a strong accent occurs every one basic beat cycle, it is a "one beat, one eye" pattern; if it occurs every three basic beat cycles, it is a "one beat, three eyes" pattern; if there is no pattern, it is a "free" pattern. Local peaks in the short-time energy envelope curve refer to those points on the curve that are higher than the values ​​of the adjacent points, i.e., peaks. Local peaks usually correspond to accents, strong beats, or emotional climaxes in singing. The basic beat cycle is the most frequently occurring and stable beat duration in a aria, usually corresponding to the time length of the "beat" or "eye" in the rhythmic pattern, and is measured in milliseconds. It is a benchmark time unit that provides a reference for determining the rhythmic pattern type and calculating the rhythmic elasticity tension. Because in traditional Chinese opera, the accents usually fall on the beat, and the beats occur periodically, the mode of the accent interval is the basic beat cycle. The beat type is the rhythmic organization pattern of traditional Chinese opera singing, defining the cyclical pattern of strong beats (beats) and weak beats (eyes). Common beat types include one beat and one eye (equivalent to 2 / 4 time, alternating between strong and weak beats), one beat and three eyes (equivalent to 4 / 4 time, strong, weak, secondary strong, weak), and free rhythm (no fixed cycle). The beat type describes the overall rhythmic framework of a aria and is an important component of the melodic parameters.

[0043] Step S402: Combine the upper limit of the vocal range, lower limit of the vocal range, commonly used registers, mode type, interval distribution histogram, and plate type into a vocal grid layer parameter vector. The vocal grid layer parameter vector is a fixed-length list of values ​​that encapsulates the structural features of opera singing that share stable commonalities across different genres and schools. It is used for genre identification, style classification, and as a skeleton constraint during AI generation. The mode type and plate type can be represented using numerical encoding. By constructing an empty list, the vocal grid layer parameters extracted in step S401 are sequentially appended to this list, and this list is output as the final vocal grid layer parameter vector.

[0044] In this embodiment, the process of extracting the cavity layer parameter vector includes: Step S403: Based on the time-stamped sequence and vocal audio data, obtain the fundamental frequency change curve over time. Within each musical pattern, divide the fundamental frequency change curve into a starting segment, a stable segment, and an ending segment. If the fundamental frequency in the starting segment continuously rises or falls by more than 0.5 semitones and the rate of change is greater than a threshold, it is determined to be a glissando. Record the glissando direction, amplitude, and duration to obtain a glissando list. If there is a reverse glissando in the stable segment after the glissando, it is marked as a return glissando. Otherwise, there are no glissandos or return glissandos, and no entries are added to the glissando list. The glissando list records the set of all detected glissando events. Each item in the list corresponds to a glissando, including the musical pattern index where the glissando is located, the glissando direction (rising or falling), the glissando amplitude (in semitones), and the glissando duration, quantifying the embellishment process of sliding from one note to another in singing. The fundamental frequency variation curve is formed by connecting the fundamental frequency values ​​extracted frame by frame in step S401 according to the time sequence of the frames. It provides the most direct raw data for subsequent detection of embellishment techniques such as glissando, vibrato, grace notes, and vocal embellishment.

[0045] Within the stable segment, the amplitude and frequency of the fundamental frequency fluctuations are detected. The difference between adjacent peaks and troughs is calculated as the fluctuation amplitude. If the difference is stable between 0.3 and 1.5 semitones and the fluctuation frequency is between 4 and 8 Hz, it is identified as a vibrato. The average depth, average frequency, and onset delay of the vibrato are recorded to obtain a vibrato list. Otherwise, it is determined that there is no vibrato, and no entry is added to the vibrato list. The vibrato list records the set of all vibrato events. Each entry includes the pitch pattern index of the vibrato, the average depth (fluctuation amplitude, in semitones), the average frequency (number of fluctuations per second), and the onset delay (how long after the start of the stable segment the vibrato appears), quantifying the periodic micro-fluctuations of pitch during singing. The process involves extracting segments of the fundamental frequency curve within a stable segment, identifying all the peaks and troughs of the fundamental frequency within that segment, and calculating the difference between the peak and trough values ​​for each pair of adjacent peaks and troughs to obtain the amplitude of a single fluctuation. The times of all peak occurrences are recorded, and the reciprocal of the time interval between adjacent peaks is calculated to obtain the frequency of each fluctuation. The average amplitude and frequency of all fluctuations are then averaged to obtain the average fluctuation amplitude and average fluctuation frequency, which are used for evaluation. The average depth of the vibrato reflects the intensity of the vibrato fluctuations; the greater the depth, the more pronounced the vibrato sounds. The average depth is obtained by summing all detected single fluctuation amplitudes and dividing by the number of fluctuations. The average frequency of the vibrato reflects the speed of the vibrato; the higher the frequency, the more rapid the vibrato. The average frequency is obtained by summing all detected fluctuation frequencies and dividing by the number of fluctuations. The onset delay reflects the time required for the singer to enter a stable vibrato state; the shorter the delay, the faster the vibrato begins, describing the vocal onset characteristics of the vibrato. By finding the position where the first fluctuation amplitude reaches half of the average depth within the stable segment, the time difference between this position and the start of the stable segment is the initial delay.

[0046] Before the start or end of a musical pattern, if a fundamental frequency segment exists with a detection duration less than the fourth time threshold of 80ms and an energy lower than 50% of the third energy ratio of the tonic, it is identified as an appoggiatura. The appoggiatura type is recorded as either a preceding or following appoggiatura and the interval distance, resulting in an appoggiatura list. Otherwise, it is determined that no appoggiatura exists, and no entry is added to the appoggiatura list. The appoggiatura list records the set of all appoggiatura events. Each item includes the appoggiatura type (preceding or following), the interval distance (the number of semitones between the appoggiatura and the tonic), and the duration, quantifying the short auxiliary sounds used as decoration. The appoggiatura type includes preceding and following appoggiaturas, obtained from the appoggiatura's temporal position relative to the tonic musical pattern: if the appoggiatura segment is immediately before the start of the tonic musical pattern, it is identified as a preceding appoggiatura; if it is immediately after the end of the tonic musical pattern, it is identified as a following appoggiatura. The interval distance of the appoggiatura is obtained from the difference between the fundamental frequency of the appoggiatura segment and the fundamental frequency of the stable segment of the tonic musical pattern. This difference is calculated in semitones and the absolute value is taken. The appoggiatura type is used to characterize whether an appoggiatura appears before or after the main note, and the interval distance quantifies the pitch difference between the appoggiatura and the main note. The start and end of the pattern are determined by detecting the fundamental frequency change rate, i.e., the method for determining the stable segment of the pattern in step S302. The fourth time threshold is the value used to distinguish between short appoggiaturas and normal short notes when judging appoggiaturas; in this example, it is 80 milliseconds. The third energy ratio is used to define the lower limit of the energy ratio weaker than the main note when judging appoggiaturas; in this example, it is 50% of the main note's energy. The fourth time threshold and the third energy ratio are based on the typical duration and energy statistics of ornaments (such as appoggiaturas and passing tones) in acoustic phonetics. The general setting range is that the fourth time threshold is between 50 and 120 milliseconds, and the third energy ratio is between 30% and 60%.

[0047] The actual duration of the current melody pattern is obtained and its ratio to the basic beat cycle is calculated. Based on this ratio, a sustained note is identified. If the ratio is greater than 1.5, it is considered a sustained note, and the duration ratio is recorded. The fundamental frequency within the sustained note segment is fitted to a straight line, and the slope (positive or negative) is used to determine the direction (ascending, descending, or flat), resulting in a sustained note list. Conversely, if the slope is less than 1.5, it is considered that no sustained note exists, and no entry is added to the sustained note list. The sustained note list records the set of all sustained note events. Each entry includes the melody pattern index, the duration ratio (actual duration divided by the basic beat cycle), and the direction (ascending, descending, or flat), quantifying the lengthening and extension of the melody on a single syllable. Specifically, when obtaining the actual duration of the current melody pattern, the record entry for the current melody pattern is found from the time stamp sequence generated in step S303, and the end and start times are directly read from this entry. In the third step, the end time is subtracted from the start time, and the resulting time difference is the actual duration of the melody pattern, in milliseconds.

[0048] Step S404 combines the glissando list, vibrato list, appoggiatura list, and sustained note list into an embellishment layer parameter vector. The embellishment layer parameter vector is a variable-length, structured list of values ​​that encapsulates the decorative features of traditional Chinese opera singing that reflect the individual singing techniques of the performers. It is used to capture subtle differences when the same aria is sung by different actors, providing key variation features for genre identification, performance level assessment, and personalized AI synthesis. An empty list is created in memory as the embellishment layer parameter vector, and all entries from all lists detected in step S403 are sequentially appended to the vector. Since the length of this vector is variable, depending on the number of embellishment events, a length field is added to the vector header during final packaging, or the vector is stored and transmitted as an independent substructure in a unified parameter model, rather than simply concatenating it into a fixed-length vector.

[0049] In this embodiment, the entire parameter extraction process from steps S401 to S404 can be implemented in a general scientific computing and digital signal processing software environment. For example, the Python programming language can be used, combined with the NumPy library for efficient numerical array operations, and signal processing modules in the SciPy library (such as scipy.signal) can be used for peak detection, autocorrelation calculation, etc., and the Librosa library (a dedicated audio analysis library) can be used for fundamental frequency extraction (such as the librosa.pyin method) and frame-level energy calculation. For the generation and segmentation logic of the time-stamped sequence, a custom Python function can be written to perform threshold judgment and boundary detection based on the energy envelope and fundamental frequency change rate array. For the judgment of glissando and vibrato, a function can be written to loop through the fundamental frequency segment within each tone pattern, use scipy.signal.find_peaks to find peaks and troughs, and calculate amplitude and frequency. All extracted parameters are ultimately stored as Python dictionaries or PandasDataFrame structures for easy concatenation later. The entire process can be encapsulated in a separate class or function, with inputs including vocal audio data, configuration vectors, and time stamp sequences, and outputs including parameter vectors for the vocal grid layer and embellishment layer.

[0050] In this embodiment, the steps for extracting dynamic statistical features from the emotional tension curve to obtain a dynamic parameter vector include: Step S501: For each analysis frame, the current fundamental frequency is normalized relative to the vocal range. The absolute value of the normalized value deviating from 0.5 is calculated as the pitch tension value, resulting in a pitch tension sequence. The generation process of the pitch tension sequence begins by obtaining the vocal range of the aria from step S401, i.e., the lowest and highest fundamental frequencies. For the fundamental frequency value of each analysis frame, its position relative to the vocal range is calculated: the lowest fundamental frequency is subtracted from the frame's fundamental frequency value, and then divided by the difference between the highest and lowest fundamental frequencies, resulting in a normalized value between 0 and 1. Then, the absolute value of this normalized value deviating from 0.5 (because 0.5 represents the center of the vocal range) is calculated to obtain the pitch tension value. Finally, the pitch tension values ​​of all frames are concatenated in chronological order to form the pitch tension sequence. The pitch tension sequence is a numerical sequence that changes over time; each value represents the degree of deviation of the sung pitch from the overall vocal range at that moment. The further the deviation from the center, the greater the tension value. The pitch tension sequence includes a tension value corresponding to each analysis frame, typically ranging from 0 to 1, quantifying the psychological tension caused by pitch rises and falls during the melody.

[0051] For each analysis frame, the dynamic tension sequence is obtained by calculating the normalized energy and the rate of energy change of adjacent frames, and then weighting them with the weight values ​​in the configuration vector. First, by scanning the entire vocal audio data, the minimum short-time energy value in all analysis frames is found. and maximum value For the energy of the current frame Calculate normalized energy , Then, the normalized energy value of the current frame and the normalized energy value of the previous frame are obtained, and the difference between the two is calculated. The absolute value of this difference is taken to obtain the energy change rate of adjacent frames. Next, the force-tension related weight values ​​are read from the configuration vector. Here, it should be the configuration weights of the two factors inside the force tensor, not the parameter set weights. For each analysis frame, its normalized energy value and the energy change rate of adjacent frames are obtained, and a weighted sum is calculated according to the configuration weights. For example, if the weight of normalized energy is 0.6 and the weight of energy change rate is 0.4, then the force-tension value = 0.6 × normalized energy + 0.4 × energy change rate. Finally, the force-tension values ​​of all frames are concatenated in chronological order to form a force-tension sequence.

[0052] The normalized energy is a value between 0 and 1 obtained by linearly mapping the short-time energy value of the current frame to the interval defined by the minimum and maximum energy of the entire singing segment. It is used to eliminate the absolute energy differences between different singing segments caused by variations in recording level and vocal loudness, making the calculation of dynamic tension comparable across different singing segments. The energy change rate between adjacent frames indicates the degree of change in the normalized energy value between the current and previous frames, reflecting the magnitude of energy change from the previous moment to the current moment and capturing sudden changes in dynamics during singing. The dynamic tension sequence is a time-varying numerical sequence, where each value combines the current volume level and the rate of volume change, representing the strength of the singing and the psychological tension caused by its changes. It includes a tension value corresponding to each analysis frame, quantifying the emotional driving force generated by loudness fluctuations during singing.

[0053] For each analysis frame, the instantaneous beat period is calculated in real time by detecting local peaks. Based on the instantaneous beat period and the basic beat period, the rhythmic elasticity tension value is calculated to obtain the rhythmic tension sequence. First, for each analysis frame, the peak in the short-time energy envelope curve is obtained as the local peak. The time interval between the current local peak and the previous local peak is calculated to represent the instantaneous beat period. Then, the basic beat period is extracted from step S401. The difference between the instantaneous beat period and the basic beat period is divided by the basic beat period to calculate the relative offset ratio. The absolute value of the relative offset ratio is taken to obtain the rhythmic elasticity tension value. Finally, since the rhythmic elasticity tension value is for the peak interval, it needs to be distributed to each analysis frame within that time interval. That is, all frames within that interval share the same rhythmic elasticity tension value. Finally, the tension values ​​of all frames are arranged in chronological order to obtain the rhythmic tension sequence.

[0054] The rhythmic elasticity tension value quantifies the degree of deviation of the current instantaneous beat cycle from the basic beat cycle of the singing segment. The greater the deviation, the greater the rhythmic elasticity tension value, representing a more unstable or elastic rhythm. It captures the emotional tension generated by the free expansion and contraction of the rhythm, the rushing or dragging of the beat during singing. The rhythmic tension sequence is a numerical sequence that changes over time. Each value represents the psychological tension generated by the rhythm deviating from the baseline beat at that moment. It includes a rhythmic elasticity tension value corresponding to each analysis frame (or the interval corresponding to each local peak), quantifying and deconstructing the emotional fluctuations brought about by the free handling of rhythm in the singing.

[0055] Step S502: Align the pitch tension sequence, dynamics tension sequence, and rhythm tension sequence by time. Read the fusion weights of pitch tension, dynamics tension, and rhythm tension from the configuration vector. These weights are the same as those assigned to the three parameter sets in step S202, used to control their proportion in emotional tension. For each moment, multiply the three tension values ​​by the corresponding weights in the configuration vector and sum them. Emotional tension value = (pitch tension weight × pitch tension value) + (dynamics tension weight × dynamics tension value) + (rhythm tension weight × rhythm tension value). Apply soft saturation mapping to the calculated emotional tension value, for example, using a hyperbolic tangent function or a sigmoid function, limiting its output range to between 0 and 1 to avoid extreme values. Finally, plot all points on a plane with the center time point of each analysis frame as the x-axis and the calculated emotional tension value as the y-axis, connecting the emotional tension values ​​of all moments to form an emotional tension curve. The emotional tension value represents the comprehensive emotional tension generated by the combined effect of the tensions of pitch, dynamics, and rhythm at a specific moment. The emotional tension curve includes the emotional tension values ​​at all moments from the beginning to the end of the singing segment. It intuitively and quantitatively displays the complete narrative arc of emotional fluctuations throughout the entire performance, clearly showing the stages of emotional establishment, climax, maintenance, and release.

[0056] Step S503: Extract the maximum, minimum, average, standard deviation, average slope of the rising segment, and area under the curve from the emotional tension curve, and combine them into a dynamic parameter vector. The dynamic parameter vector is a fixed-length list of values ​​that extracts macroscopic statistical features from the emotional tension curve to summarize the overall emotional dynamics of the entire singing passage. Specifically, the maximum value of the emotional tension curve represents the tension value at the most tense moment of the entire piece, the minimum value represents the tension value at the most relaxed moment, the average value represents the average tension level, the standard deviation represents the intensity of emotional fluctuations, the average slope of the rising segment represents the average speed of emotional rise, and the area under the curve represents the time integral of emotional tension, i.e., the total emotional load. This further compresses the one-dimensional, variable-length emotional tension curve into several key statistical quantities, facilitating rapid emotional classification, emotional comparison between different singing segments, or serving as input features for machine learning models.

[0057] In this embodiment, the process of obtaining the unified parameter model includes: Step S601: The cavity layer parameter vector, the embellishment layer parameter vector, and the dynamic parameter vector are each packaged into a fixed-length numerical list, and then concatenated sequentially into a one-dimensional statistical parameter vector. The statistical parameter vector integrates the features describing different dimensions of opera singing into a compact, fixed-length mathematical representation.

[0058] Step S602: Adjust the list of values ​​in the statistical parameter vector based on the downstream task labels; determine whether to include the emotional tension curve based on the downstream task labels, and finally obtain the unified parameter model. When adjusting based on the downstream task labels: If the downstream task is categorized or compared and the configuration vector requires compression, then the mean and variance of the glissando amplitude, vibrato frequency, and phrasing ratio in the embellishment layer parameter vector are calculated, and the original list is replaced to compress the data volume. For scenarios where the downstream task is categorized or compared, the focus is on the overall statistical distribution of embellishment features rather than the details of each event. Using the mean and variance can effectively compress a variable-length list into a fixed-length statistical quantity, while preserving the central tendency and dispersion information of the data, significantly reducing the dimensionality of the feature vector, reducing data redundancy, and improving the generalization performance of the classifier. When performing specific mean and variance calculations, firstly, the glissando list is extracted from the original embellishment layer parameter vector. The list is traversed to extract the amplitude value of each glissando, forming a one-dimensional array. All amplitude values ​​are summed and divided by the total number of glissandos to obtain the mean of the glissando amplitude. The mean is subtracted from each amplitude value, the square is taken, all squares are summed, and then divided by the total number of glissandos to obtain the variance of the glissando amplitude. Then, following the same method of calculating the mean and variance of the amplitude values ​​for the glissando list, the mean frequency value for the vibrato list and the duration ratio for the sustain list are calculated sequentially, yielding the mean and variance of the vibrato frequency and the mean and variance of the sustain ratio, respectively. Finally, these three means and three variances are used to replace the original glissando list, vibrato list, and sustain list.

[0059] If the downstream task tag is "teaching," then the pre-stored standard template parameters corresponding to the genre and style tags are retrieved. The difference between the current extracted value and the standard value is calculated parameter by parameter, resulting in a deviation vector. First, ensure that the parameter list of the current extracted value and the parameter list of the standard value have the exact same order and length. Create an empty list as the deviation vector, and extract each parameter value from the current extracted value and the corresponding parameter value from the standard value, from left to right. Then, calculate the difference between the standard value and the current extracted value, append this difference to the deviation vector, and repeat the difference calculation process until all parameters have been processed. This point-by-point difference calculation provides the most direct and transparent way to quantify differences, generating a multi-dimensional deviation report that clearly indicates the gap between the student's performance and the standard singing in specific aspects such as vocal range, pitch accuracy, rhythm, and vocal embellishment techniques. The parameter-by-parameter parameter includes all parameters marked as comparable in the unified parameter model, specifically covering parameters at the vocal register level (such as upper and lower bounds of the vocal range, boundaries of commonly used vocal registers, and the proportion of each interval), parameters at the embellishment level (such as the average amplitude of glissando, the average frequency of vibrato, and the average proportion of sustained notes), and dynamic parameters (such as the maximum and average values ​​of emotional tension, and the slope of the rising segment). The current extracted value is obtained from the statistical parameter vector generated in step S601, and the standard value is obtained from the pre-stored template library by querying based on the currently input genre and style tags.

[0060] A unified parameter model is created by combining a one-dimensional statistical parameter vector, an emotional tension curve, and a deviation vector. This unified parameter model is a structured data object, providing a comprehensive digital description of a segment of opera singing. It integrates abstract features from the statistical level with dynamic curves from the time-series level, including a mandatory one-dimensional statistical parameter vector, an optional complete emotional tension curve, and an optional deviation vector. As a universal and scalable data exchange format, the unified parameter model provides upper-layer applications with on-demand feature subsets, avoiding redundant calculations and enabling analysis once and reuse across multiple scenarios.

[0061] The emotional tension curve and deviation vector are optional components, and their inclusion depends on the downstream task labels. The emotional tension curve contains a large amount of data; for tasks requiring only statistical features, such as classification and comparison, it is redundant information, and including it would waste storage and transmission bandwidth. However, it is essential for tasks requiring temporal details, such as visualizing teaching presentations and AI-generated frame-by-frame control. The deviation vector is only meaningful in teaching scenarios where comparison with a standard template is required; in other scenarios, the concept of a standard value does not exist, therefore it should not be included.

[0062] The pre-stored opera genre tags include a standard template library. This library pre-stores a set of standard parameter vectors for each opera genre (such as Qinqiang, Peking Opera, Henan Opera, etc.) and each school (such as Dan (female role), Sheng (male role), Mei School, etc.). These standard parameter vectors specifically include typical values ​​for the opera genre or school at the vocal structure, embellishment, and dynamics levels, such as standard vocal range, typical vibrato depth, and standard emotional tension curve shape. These values ​​are derived from parametric analysis of classic arias by numerous renowned opera performers, extracting statistically significant typical values. These values, after review and confirmation by opera music experts, serve as a golden template for comparison in teaching tasks. They are used to calculate the deviation between student singing and standard singing, thereby providing quantitative and targeted teaching feedback.

[0063] In this embodiment, the implementation steps for selecting and generating corresponding output feature vectors from the unified parameter model based on downstream task labels include: Step S701: Read downstream task labels, which include classification labels, comparison labels, teaching labels, and AI-generated labels. Downstream task labels are directly obtained from the configuration vector generated in step S200. These labels are initially input by the user when initiating the analysis request and are then solidified into the configuration vector through the configuration mapping table. Downstream task labels are categorized into classification labels, comparison labels, teaching labels, and AI-generated labels according to the application scenario. Classification labels represent the goal of this analysis: to classify the current singing segment into a predefined category (e.g., determining which opera genre or school the singing segment belongs to), requiring the output of features most suitable for the classification algorithm. Comparison labels represent the goal of this analysis: to compare the similarity between the current singing segment and another singing segment, requiring the output of features convenient for calculating distance or similarity. Teaching labels represent the goal of this analysis: to assist in opera teaching by comparing students' singing with standard templates, requiring the output of detailed feedback including deviation details and visualization curves. AI-generated labels represent the goal of this analysis: to provide control conditions for the artificial intelligence generation model, requiring the output of parameters that guide the synthesizer to generate singing with specific styles and emotions.

[0064] Step S702: Based on the downstream task labels, select the corresponding vector from the unified parameter model as the output feature vector, and send the output feature vector to the downstream application through the interface; when generating the output feature vector through the downstream labels: If the downstream task label is classification, a one-dimensional statistical parameter vector is extracted from the unified parameter model as the output feature vector. For classification labels, the basis for generating the output feature vector is that classification algorithms typically require statistical features with fixed dimensions and stable numerical ranges, but do not require temporal details. The goal is to output a compact feature vector that can be directly input into a classifier (such as a support vector machine or a neural network classification layer).

[0065] If the downstream task label is "comparison," a one-dimensional statistical parameter vector is extracted from the unified parameter model, and the cosine similarity or Euclidean distance between the parameter vector of the current segment and the parameter vector of another segment is calculated as the output. For the comparison label, the core of the comparison task is to quantify the difference or similarity between two entities. Directly outputting the distance value or similarity score is the simplest and most effective result, allowing the caller to obtain comparable numerical results directly without having to calculate the distance between complex features themselves.

[0066] If the downstream task is tagged as "teaching," then the deviation vector and emotional tension curve are extracted from the unified parameter model, and a textual feedback report is generated that includes numerical descriptions of the deviation and curve comparisons. For the "teaching" tag, the teaching feedback needs to specify where the singing went wrong, by how much, and provide an intuitive visual comparison, generating a teaching report that includes quantitative deviation values ​​and an emotional tension curve comparison chart to help learners intuitively understand the problem.

[0067] If the downstream task label is AI-generated, then a one-dimensional statistical parameter vector and emotional tension curve are extracted from the unified parameter model as conditional parameters for output. Modern AI-generated labels (such as vocoders based on diffusion or GANs) often require multi-layered conditioning information, including global statistical features and local temporal features, providing a rich set of conditional parameters to guide the generation model in synthesizing opera singing styles with specific characteristics and emotional fluctuations.

[0068] The process of calculating the cosine similarity or Euclidean distance between the current vocal segment parameter vector and another vocal segment parameter vector as output includes the following steps: First, extract a one-dimensional statistical parameter vector from the unified parameter model of the current aria, denoted as vector A. Extract the statistical parameter vector of the other aria using the same method, denoted as vector B. Ensure that vectors A and B have the same dimensions and that all parameters are aligned. Then, choose either cosine similarity or Euclidean distance as the output. If you choose to output cosine similarity, then calculate the dot product of vector A and vector B, which is to multiply the values ​​at corresponding positions in the two vectors, and then add all the product results. Calculate the magnitude of vector A and vector B separately, which is the square root of the sum of the squares of all the values ​​of each vector. The cosine similarity is equal to the dot product divided by the product of the magnitudes of vector A and vector B.

[0069] If you choose to output Euclidean distance, then calculate the difference vector between vector A and vector B, which is to subtract the corresponding values, square each value in the difference vector, and then add all the squared values ​​together. The Euclidean distance is equal to the square root of the sum.

[0070] Cosine similarity measures the directional consistency of two vectors and is insensitive to the absolute magnitude of the vectors. It is suitable for comparing the similarity in shape between vocal segments of different loudness or energy levels. Euclidean distance measures the absolute distance between two vectors in space and is sensitive to magnitude. It is suitable for comparing the closeness of overall features and simplifies the complex, high-dimensional relationship of vocal features into an intuitive, sortable single value. It can be used to realize functions such as vocal segment retrieval, style clustering, or automatic scoring.

[0071] like Fig. 3 As shown, the intelligent parameterized extraction and representation system for opera singing features includes: The audio acquisition module is used to acquire the original opera singing audio and preprocess it to obtain singing audio data; The parameter configuration module is used to obtain the input genre tags, style tags, and downstream task tags, and generate configuration vectors by querying the pre-stored configuration mapping table; The vocal structure segmentation module is used to perform structural segmentation on vocal audio data, detect breathing points, cavity boundaries and melodic boundaries, and obtain a time-stamped sequence containing musical phrases, cavities and melodic patterns. The feature extraction module is used to extract the cavity grid layer parameter vector and the embellishment layer parameter vector from the audio segments corresponding to the time stamp sequence in the vocal audio data based on the configuration vector; The dynamics parameter module is used to generate an emotional tension curve based on the configuration vector, vocal audio data, time stamp sequence, and cavity grid parameter vector, and to extract dynamic statistical features from the emotional tension curve to obtain the dynamics parameter vector. The parameter fusion module is used to structurally fuse the cavity layer parameter vector, the lubrication layer parameter vector, the dynamic parameter vector, and the emotional tension curve to obtain a unified parameter model; The adaptive output module is used to select and generate corresponding output feature vectors from the unified parameter model based on the downstream task labels.

[0072] In this embodiment, a configuration mapping table and rule engine containing genre tags, style tags, and downstream task tags are constructed. Before the analysis begins, a configuration vector containing parameter set switches, weight values, and output granularity information is automatically generated. Based on the downstream tags such as classification, comparison, teaching, and AI generation, the weight values ​​and output granularity of the set parameters are automatically adjusted to achieve adaptive information configuration. With the configuration vector as a constraint, corresponding feature extraction is performed, ensuring that the content after feature extraction corresponds one-to-one with the needs of the downstream tags. This solves the inherent problems of one-size-fits-all and mixed feature extraction in the prior art, allowing the same method to adapt to different application scenarios without modifying the algorithm. Dynamic parameter adjustment and feature extraction can be performed for different genres, styles, and tasks, effectively improving the accuracy of parameter configuration. This is superior to the comparison scheme using a fixed parameter set, achieving applicability and flexibility across genres, styles, and multiple tasks, and increasing the accuracy of feature extraction and analysis.

[0073] This method decomposes and integrates the features of traditional Chinese opera singing into three levels. The extraction of parameter vectors from the vocal structure layer reflects the stylized common structures of different opera genres and schools, exhibiting high stability. The extraction of parameter vectors from the embellishment layer reflects the individualized singing techniques of actors, exhibiting high variability. The calculation of emotional tension curves extracts parameters from the dynamics layer, quantifying abstract emotional fluctuations into a computable time series. These three layers of parameters are structurally fused to form a unified parameter model. The model can be output as a complete model or a subset depending on the downstream task. This allows traditional Chinese opera musicology research to focus on vocal structure layer parameters, genre identification to combine vocal structure and embellishment layers, and emotion calculation and AI generation to rely on dynamics layer parameters. This achieves a leap from single-feature to multi-dimensional decoupled integration, enabling targeted parameter generation strategies for different downstream tasks. It improves the accuracy of opera genre classification, the precision of teaching feedback, and the depth of understanding of emotion analysis, representing significant progress in feature representation completeness and application versatility.

[0074] This application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memories store computer-readable code, which, when executed by the one or more processors, can perform the intelligent parameterization extraction and characterization method and system for opera singing features as described above.

[0075] The methods and systems according to the embodiments of this application can also be implemented using the architecture of the electronic device shown in this application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. The storage device in the electronic device, such as a ROM or hard disk, may store the intelligent parameterization extraction and characterization method and system for opera singing features provided in this application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in this application is merely exemplary; when implementing different devices, one or more components in the electronic device shown in this application may be omitted according to actual needs.

[0076] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising a reference structure" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0077] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for intelligent parameterized extraction and representation of opera singing features, characterized in that, The method includes: The original opera singing audio is acquired and preprocessed to obtain the singing audio data; Obtain the input genre tags, style tags, and downstream task tags, and generate a configuration vector by querying a pre-stored configuration mapping table; The vocal audio data is structurally segmented to detect breathing points, cavity boundaries, and melodic boundaries, resulting in a time-stamped sequence containing musical phrases, cavities, and melodic patterns. Based on the configuration vector, extract the cavity grid layer parameter vector and the embellishment layer parameter vector from the audio segments corresponding to the time stamp sequence in the vocal audio data; Based on the configuration vector, vocal audio data, time stamp sequence, and cavity layer parameter vector, the pitch tension sequence, dynamic tension sequence, and rhythmic elasticity tension sequence are calculated and synthesized into an emotional tension curve. Dynamic statistical features are extracted from the emotional tension curve to obtain a dynamic parameter vector. The cavity layer parameter vector, the cavity lubrication layer parameter vector, the dynamic parameter vector, and the emotional tension curve are structurally fused to obtain a unified parameter model; Based on the downstream task labels, the corresponding output feature vectors are selected from the unified parameter model to generate.

2. The intelligent parameterized extraction and characterization method for opera singing features according to claim 1, characterized in that, The process of acquiring and preprocessing the original opera singing audio to obtain singing audio data includes: The acquired original opera singing audio is converted into a mono signal with a uniform sampling rate; Calculate the energy of each segment within a preset time window, and cut off segments whose energy is lower than the first energy ratio of the maximum energy and whose duration exceeds the first time threshold; The amplitude distribution of the cut segments is statistically analyzed, and the amplitude threshold is set to the first multiple of the maximum amplitude. Samples in the audio that are below the amplitude threshold are set to zero. The maximum amplitude of the entire audio is adjusted to obtain the vocal audio data.

3. The intelligent parameterized extraction and characterization method for opera singing features according to claim 1, characterized in that, The process of obtaining the input genre tags, style tags, and downstream task tags, and generating a configuration vector by querying a pre-stored configuration mapping table, includes: The pre-stored configuration mapping table is searched using the combination of obtained genre tags, genre tags, and downstream task tags as an index. If an exact match exists, the corresponding configuration vector is read directly; if no exact match exists, the closest configuration is matched according to the priority order of genre, followed by style, and finally downstream task, and a configuration vector is generated according to preset configuration rules. The configuration vector includes a parameter set switch list, a weight value list, and an output granularity. The parameter set switch list includes a cavity grid parameter set switch, a embellishment parameter set switch, and a dynamic parameter set switch, and each parameter set switch is assigned a weight value. The output granularity includes sentence level, cavity level, and tone pattern level.

4. The intelligent parameterized extraction and characterization method for opera singing features according to claim 1, characterized in that, The process of structurally segmenting the vocal audio data, detecting breath points, phrasal boundaries, and melodic boundaries, yields a time-stamped sequence containing musical phrases, phrasal sections, and melodic patterns, including: The vocal audio data is divided into a frame sequence with a specified frame length and a specified frame shift. The sum of squares of the sampling points in each frame is calculated to obtain the short-time energy envelope curve. In the short-time energy envelope curve, the position where the energy drops sharply beyond the second energy ratio and the duration exceeds the second time threshold is marked as a breathing point. The interval between adjacent breathing points is a musical phrase, and the start and end times of each musical phrase are recorded. Within each musical phrase, the fundamental frequency change rate between adjacent frames is calculated. The position where the absolute value of the fundamental frequency change rate exceeds the first preset multiple of the average change rate and is accompanied by a decrease in energy is marked as the cavity boundary. The start and end times of each cavity are recorded. When the output granularity in the configuration vector is at the tone pattern level, the alternation point between the fundamental frequency stable segment and the fundamental frequency changing segment is detected within each cavity segment, the alternation point is defined as the tone pattern boundary, and the start and end times of each tone pattern are recorded. The start and end times of each musical phrase, each melody, and each musical pattern are combined into a time-stamped sequence.

5. The intelligent parameterized extraction and characterization method for opera singing features according to claim 4, characterized in that, The process of extracting the cavity lattice layer parameter vector includes: For each musical phrase, extract the fundamental frequency of all frames, find the lowest and highest fundamental frequencies as the lower and upper limits of the pitch range, and determine the commonly used pitch range; Map all fundamental frequencies to pitch class numbers based on the tonic, generate a pitch class histogram by counting the frequency of each pitch class, and determine the mode type based on the peak patterns of the pitch class histogram: Within a musical phrase, the fundamental frequency difference between adjacent musical patterns is calculated sequentially, and the distribution of all differences is statistically analyzed to generate an interval distribution histogram. Extracting the local peaks of the short-time energy envelope curve as accent positions, and statistically analyzing the intervals with the highest frequency among adjacent accent intervals as the basic beat cycle, the type of rhythm is determined based on the pattern of strong accent occurrences. The upper limit of the pitch range, the lower limit of the pitch range, the commonly used pitch range, the mode type, the interval distribution histogram, and the plate type are combined into a cavity grid layer parameter vector.

6. The intelligent parameterized extraction and characterization method for opera singing features according to claim 5, characterized in that, The process of extracting the cavity layer parameter vector includes: Based on time-stamped sequences and vocal audio data, the curve of fundamental frequency change over time is obtained, and when glissando is present, a list of glissando notes is obtained. Within the stable range, the amplitude and frequency of the fundamental frequency fluctuations are detected, and a list of vibratos is obtained when a vibrato is detected. If, before the start or end of a musical pattern, there exists a fundamental frequency segment whose detection duration is less than the fourth time threshold and whose energy is lower than the third energy ratio of the main tone energy, it is determined to be an appoggiatura and an appoggiatura list is obtained. Obtain the actual duration of the current melody and calculate the ratio to the basic beat cycle. Determine the duration of the sustained notes based on the ratio. If sustained notes exist, obtain a list of sustained notes. Combine the glissando list, vibrato list, appoggiatura list, and embellishment list into an embellishment layer parameter vector.

7. The intelligent parameterized extraction and characterization method for opera singing features according to claim 1, characterized in that, The process of extracting dynamic statistical features from the emotional tension curve to obtain a dynamic parameter vector includes: For each analysis frame, the current fundamental frequency is normalized relative to the pitch range, and the absolute value of the normalized value deviating from 0.5 is calculated as the pitch tension value to obtain the pitch tension sequence. For each analysis frame, the force tension sequence is obtained by calculating the normalized energy and the energy change rate of adjacent frames, and then weighting and summing them with the weight values ​​in the configuration vector. For each analysis frame, the instantaneous beat cycle is calculated in real time by detecting local peaks, and the rhythmic elasticity tension value is calculated based on the instantaneous beat cycle and the basic beat cycle to obtain the rhythmic tension sequence; The pitch tension sequence, dynamic tension sequence, and rhythm tension sequence are aligned by time. At each moment, the three tension values ​​are multiplied by the corresponding weights in the configuration vector and summed. The emotional tension value is obtained by soft saturation mapping. The emotional tension values ​​at all moments are connected to form an emotional tension curve. The maximum, minimum, average, standard deviation, average slope of the rising segment, and area under the curve are extracted from the emotional tension curve and combined into a dynamic parameter vector.

8. The intelligent parameterized extraction and characterization method for opera singing features according to claim 7, characterized in that, The process of obtaining the unified parameter model includes: The cavity layer parameter vector, the cavity lubrication layer parameter vector, and the dynamic parameter vector are each packaged into a fixed-length numerical list, and then concatenated in order into a one-dimensional statistical parameter vector. The list of values ​​in the statistical parameter vector is adjusted based on the downstream task labels; the inclusion of the emotional tension curve is determined based on the downstream task labels, and finally a unified parameter model is obtained.

9. The intelligent parameterized extraction and characterization method for opera singing features according to claim 1, characterized in that, The step of selecting and generating corresponding output feature vectors from the unified parameter model based on the downstream task labels includes: Read the downstream task tags, which include classification tags, comparison tags, teaching tags, and AI-generated tags; Based on the downstream task labels, the corresponding vectors are selected from the unified parameter model as output feature vectors; The output feature vector is sent to the downstream application via an interface.

10. A system for intelligent parameterized extraction and representation of opera singing characteristics, characterized in that, The system includes: The audio acquisition module is used to acquire the original opera singing audio and preprocess it to obtain singing audio data; The parameter configuration module is used to obtain the input genre tags, style tags, and downstream task tags, and generate configuration vectors by querying the pre-stored configuration mapping table; The vocal structure segmentation module is used to perform structural segmentation on the vocal audio data, detect breathing points, cavity boundaries and melodic boundaries, and obtain a time-stamped sequence containing musical phrases, cavities and melodic patterns. The feature extraction module is used to extract the cavity grid layer parameter vector and the embellishment layer parameter vector from the audio segments corresponding to the time stamp sequence in the singing audio data based on the configuration vector. The dynamics parameter module is used to generate an emotional tension curve based on the configuration vector, vocal audio data, time stamp sequence, and cavity grid parameter vector, and to extract dynamic statistical features from the emotional tension curve to obtain a dynamics parameter vector. The parameter fusion module is used to structurally fuse the cavity layer parameter vector, the lubrication layer parameter vector, the dynamic parameter vector, and the emotional tension curve to obtain a unified parameter model; An adaptive output module is used to select and generate corresponding output feature vectors from a unified parameter model based on the downstream task labels.