Multi-feature fusion-based drama cavity modeling method and system
Through the multi-feature fusion opera aria modeling method, a two-stage dynamic benchmark model was constructed, which solved the problem of insufficient dynamic process modeling in the vocal preparation stage in the existing technology, realized the full-process analysis and guidance of opera aria, and improved the evaluation of predictability and artistic expression.
Patent Information
- Application Number
- CN202511111540.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies lack the ability to model and analyze the dynamic processes of the vocal preparation stage in opera aria analysis methods, resulting in deficiencies in the fundamentality and predictability of guiding information. They are unable to effectively capture the coordinated relationship between multiple dimensions such as pitch, timbre, and stability, and lack flexible evaluation criteria for artistic expression.
By adopting the multi-feature fusion method, the multi-dimensional feature vector of the audio signal is obtained to generate the dynamic dominant time series and perform phase space reconstruction, and a two-stage dynamic benchmark model is constructed, including the ideal precursor trajectory cluster in the vocal preparation stage and the benchmark artistic manifold in the vocal performance stage, to generate feedforward and feedback guidance information.
It achieves a comprehensive and fundamental analysis of the entire process of singing behavior, improves the comprehensiveness and depth of guidance information, enhances foresight and initiative, and improves the objectivity and precision of the description of singing behavior.
Smart Images

Figure CN120748432A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of acoustic signal processing and pattern recognition, and in particular to a multi-feature fusion opera aria modeling method and system. Background Art
[0002] Opera singing is a performing art that demands extremely high technical skill and artistic expression. Traditional singing training relies primarily on instruction from professional vocal teachers. However, this approach is costly, subjective, and difficult to quantify. With the advancement of computer technology, systems and methods have emerged that utilize acoustic analysis techniques to assist singing training.
[0003] However, existing singing training methods typically rely on analyzing a single or a few isolated acoustic features, such as tracking only the accuracy of pitch (fundamental frequency) or volume. This approach fragments the complex vocalization process into discrete metrics and fails to capture the inherent coordination and dynamic evolution of multiple dimensions such as pitch, timbre, and stability. A successful singing technique is essentially a specific dynamic pattern formed by the coordinated movement of multiple physiological and acoustic parameters. However, existing technologies, due to the limitations of their analytical dimensions, are unable to model and evaluate this holistic dynamic pattern.
[0004] In addition, the guidance information provided by existing technologies is mostly instant feedback. That is, the system will judge the singer's performance only after the singer makes a sound, for example, pointing out that the pitch deviates from the target value. This feedback mode is passive. It only focuses on the results that have occurred, and ignores the fact that for a note, especially a challenging high or long note, the preparation stage before the sound is the key to determining its quality. How the singer adjusts his breathing and adjusts the cavity resonance to "prepare" to make a high-quality sound is completely ignored in the existing technology. Therefore, the existing technology lacks effective modeling and forward-looking guidance mechanism for vocal preparation actions.
[0005] At the same time, in the vocal performance stage, the evaluation criteria of existing technologies are often rigid, such as setting the target pitch to a fixed frequency value or a narrow range. This approach cannot accommodate and understand the necessary artistic processing techniques in opera singing, such as vibrato, glissando, and the dynamic changes in timbre brought about by emotional expression. These techniques are acoustically manifested as characteristic dynamic fluctuations. If fixed values are used as standards, they will be misjudged as errors or instability. Therefore, the existing technology lacks a flexible benchmark model that can describe and accommodate the "correct state space" formed by artistic expression, making it impossible to effectively analyze and guide the artistry of singing.
[0006] To sum up, how to establish an overall dynamic model that can simultaneously reflect the vocal preparation stage and the vocal performance stage from the perspective of multi-dimensional feature fusion, and provide guidance that is both forward-looking and artistically inclusive, is a technical problem that needs to be solved urgently in this field. Summary of the Invention
[0007] In response to the shortcomings of the existing technology, the present invention provides a multi-feature fusion opera aria modeling method and system, which solves the problem that the existing opera aria analysis methods usually only perform static or isolated evaluations of the acoustic features during the vocalization process, lack the ability to model and analyze the entire singing behavior, especially the dynamic process of the vocalization preparation stage, resulting in the guidance information generated by them being insufficient in fundamentality and predictiveness.
[0008] To achieve the above objectives, the present invention provides a first aspect of a multi-feature fusion opera aria modeling method, which comprises the following steps: First, an audio signal to be analyzed is obtained, from which a multidimensional feature vector sequence is extracted, wherein each multidimensional feature vector contains multiple preset acoustic features for describing the singing state at a specific moment.
[0009] Subsequently, a one-dimensional dynamics-dominated time series is generated based on the multidimensional eigenvector sequence. This step can be specifically implemented using principal component analysis. The continuous multidimensional eigenvectors are formed into a data matrix, the covariance matrix of the data matrix is calculated and the eigenvalue decomposition is performed to obtain the first principal component direction. The multidimensional eigenvector sequence is projected onto the first principal component direction to generate the dynamics-dominated time series. The sequence It can be expressed by the following formula: ; in, For the moment The multidimensional feature vector of is the eigenvector corresponding to the direction of the first principal component.
[0010] Then, the phase space is reconstructed using the dynamic dominant time series to obtain the phase space trajectory that characterizes the dynamic state of the singer. This step is achieved by the time delay embedding method. Build a dimensional dynamical state vector : ; in, For the moment The dynamical state vector of For the moment The dynamics of the dominant time series, is the embedding dimension, For time delay; The phase space trajectory is composed of the continuous dynamical state vector constitute.
[0011] Then, based on the preset sample singing data, a two-stage dynamic benchmark model is constructed in phase space. The construction of this model includes: Perform note starting point detection on the sample singing data to obtain a series of note starting point timestamps.
[0012] The phase space state point corresponding to the timestamp of the note starting point is marked as the ideal starting point.
[0013] The set of phase space trajectories within the preset time window before each ideal take-off point is constructed as an ideal precursor trajectory cluster, which is used to characterize the dynamic criteria of the phonation preparation stage.
[0014] The phase space trajectory point cloud during the note duration is processed by a manifold learning algorithm (such as the Isomap algorithm) to construct a benchmark art manifold for characterizing the dynamic standards of the vocal performance stage.
[0015] Finally, based on the comparative analysis of the phase space trajectory generated by the audio signal to be analyzed and the two-stage dynamics benchmark model, targeted singing guidance information is generated. This process specifically includes two stages: In the vocalization preparation phase, the phase space trajectory generated by the user is tracked as the user's precursor trajectory, and the predicted starting point to be reached is predicted based on the user's precursor trajectory. The difference between the predicted starting point and the ideal starting point of the corresponding note is calculated to generate feedforward guidance information. The feedforward guidance information can be a feedforward guidance vector , which is calculated as follows: ; in, For the moment The feedforward guidance vector, For the The ideal starting point for each note, For the moment For the first The predicted starting point of each note, is the proportional gain coefficient, is the differential gain coefficient, is a standard mathematical operator.
[0016] During the vocalization phase, the user's current phase-space state point is obtained, the deviation between this state point and the reference art manifold is calculated, and its projection point on the reference art manifold is determined. Based on the relationship between this state point and the projection point, feedback guidance information is generated to guide the user's phase-space state point toward the reference art manifold.
[0017] A second aspect of the present invention provides a multi-feature fusion opera aria modeling system, the system comprising: A signal processing module configured to obtain an audio signal to be analyzed and extract a multidimensional feature vector reflecting a singing state from the audio signal; a dynamics modeling module, connected to the signal processing module, configured to generate a one-dimensional dynamics-dominant time series based on the multidimensional feature vector, and to perform phase space reconstruction using the dynamics-dominant time series to obtain a phase space trajectory representing the dynamic state of the singer; A benchmark model storage module, configured to store a preset two-stage dynamics benchmark model constructed in phase space, the two-stage dynamics benchmark model comprising an ideal precursor trajectory cluster and an ideal take-off point for characterizing a vocal preparation phase, and a benchmark art manifold for characterizing a vocal performance phase; An analysis and guidance module is connected to the dynamic modeling module and the benchmark model storage module, and is configured to compare and analyze the phase space trajectory generated by the dynamic modeling module with the two-stage dynamic benchmark model stored in the benchmark model storage module, and generate targeted singing guidance information accordingly.
[0018] The present invention provides a multi-feature fusion opera aria modeling method and system. It has the following beneficial effects: 1. By constructing a two-stage dynamic benchmark model consisting of an ideal precursor trajectory cluster, an ideal take-off point, and a benchmark artistic manifold, this method expands the scope of analysis from a single vocal performance stage to the entire vocal preparation process. This enables a more complete modeling of the singing behavior, breaking through the limitation of existing technologies that only analyze the produced sounds. This allows for a comprehensive and fundamental analysis of the entire singing process, enhancing the comprehensiveness and depth of the guidance information.
[0019] 2. This invention analyzes the user's precursor trajectory during the vocal preparation phase, predicts the difference between it and the ideal take-off point, and generates feedforward guidance information accordingly. This mechanism enables the system to intervene before vocal errors actually occur, preemptively correcting the singer's preparatory movements. This enhances the predictability and proactive nature of guidance information, changing the traditional approach of reactive correction after errors have occurred.
[0020] 3. By generating a dynamics-dominant time series and performing phase space reconstruction, this invention transforms multidimensional, isolated acoustic features into a phase space trajectory that holistically reflects the inherent dynamics of the singing system. This elevates the analysis dimension from static eigenvalue matching to dynamic conformance assessment of state transitions, capturing the transitional relationships between notes and improving the objectivity and sophistication of the model's description of singing behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of the method of the present invention; Figure 2 This is a system architecture diagram of the present invention.
[0022] Among them, 10, signal processing module; 20, dynamic modeling module; 30, benchmark model storage module; 40, analysis and guidance module. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0024] Example: Please see the attached Figure 1 The embodiment of the present invention provides a multi-feature fusion opera aria modeling method, comprising the following steps: S201, obtaining an audio signal to be analyzed, and extracting a multidimensional feature vector reflecting a singing state therefrom; This step forms the data foundation for all subsequent dynamic analysis, and its execution quality directly impacts the accuracy of the final model. This step specifically includes two subprocesses: signal acquisition and purification, and short-term feature extraction. This step extracts features by framing the audio stream into short frames (for example, one frame every 20 milliseconds). Algorithms in modern signal processing libraries (such as FFT and MFCC calculations) are highly optimized and can complete calculations within a single frame time window, enabling streaming processing.
[0025] In a specific embodiment, the signal acquisition is performed by recording the singer's audio in digital form through one or more microphones. To ensure that the signal can fully reflect the singing details, the acquisition parameters are set to high standards, such as the sampling rate. Set to 44100Hz or higher, quantization bit depth Set to 24 bits or higher. The original audio signal collected is represented as a discrete time series Since the original signal contains environmental noise or accompaniment music, it needs to be purified to separate the pure vocal signal. This purification process can use the Wiener filtering algorithm, which estimates the power spectral density of the signal and noise, designs an optimal filter, and filters the signal to suppress the non-human voice part.
[0026] Get pure vocal signal Then, it is framed. The signal is divided into a series of lengths of There is a short time frame of sampling points, and there is a gap of length The frame shift of sampling points is To ensure continuity between frames. For example, the frame length can be set to 40 milliseconds and the frame shift can be set to 10 milliseconds. To reduce spectral leakage, each frame signal is multiplied by a window function (such as a Hamming window) before feature extraction.
[0027] For each frame of the windowed signal, extract an N-dimensional feature vector The vector is composed of multiple acoustic features, which are used to describe the vocalization state at that moment from different dimensions. In a specific embodiment, these acoustic features include: Fundamental frequency: This represents the frequency of vocal cord vibration and directly corresponds to the pitch of the sung note. It is the basis for melodic analysis, and its stability reflects the singer's ability to control pitch.
[0028] Harmonic-to-Noise Ratio: Quantifies the energy ratio of the harmonic components to the noise components in the signal. A high HNR value generally corresponds to a clear, stable vocalization, while a low value indicates breathiness or roughness of the voice.
[0029] Spectral Centroid: This represents the center frequency of the spectral energy distribution. This characteristic is directly related to the brightness of the timbre: a higher spectral centroid corresponds to a brighter timbre, and vice versa. Its variations can reflect the singer's control of timbre when expressing emotion.
[0030] Mel-frequency cepstral coefficients: The first 13 coefficients are typically extracted. They provide a compact representation of the spectral envelope that matches the nonlinear frequency perception of the human ear, effectively characterizing timbre. They are key features for distinguishing different vowels and timbre qualities.
[0031] The individual scalar or vector features extracted above are combined into a single multidimensional feature vector in each frame For example, if we extract one fundamental frequency value, one HNR value, one spectral centroid value, and 13 MFCCs coefficients, we construct a 16-dimensional feature vector. This process is performed continuously on the entire audio signal, ultimately generating a time series of multidimensional feature vectors, which serves as input data for the next step of dynamic analysis.
[0032] S202, generating a one-dimensional dynamics-dominated time series based on the multi-dimensional eigenvector; This step aims to extract a one-dimensional signal that can best reflect the overall dynamic changes of the system from a high-dimensional feature space with correlated features. This signal simplifies the complexity of subsequent dynamic analysis while retaining the most critical dynamic information. In a specific embodiment, this step is implemented using a principal component analysis algorithm, the core of which is the application of principal component analysis (PCA). PCA training (i.e., finding the direction of the principal component) is completed offline together with the benchmark model. When applied online, for each new multidimensional feature vector, only one matrix-vector multiplication is required to generate its dominant component. This is a computationally extremely low-cost operation that can be completed instantaneously.
[0033] First, the continuous multi-dimensional feature vector sequence obtained in step S201 is organized into a data matrix If there are time frames, each feature vector is dimension, then the data matrix The dimension is . Each row of the matrix corresponds to the eigenvector at a moment.
[0034] Next, the data matrix Perform centralization, that is, subtract the mean of each column (corresponding to a feature dimension). Then, calculate the covariance matrix of the centered data matrix The diagonal elements of the covariance matrix represent the variance of each feature itself, and the off-diagonal elements represent the covariance between different features, thus capturing the linear correlation of all features.
[0035] Pair covariance matrix Perform eigenvalue decomposition to obtain a set of eigenvalues and the corresponding eigenvector These eigenvalues are arranged in descending order, i.e. Each eigenvector Represents the original An orthogonal direction in the dimensional feature space, and its corresponding eigenvalue It indicates the variance of the original data in this direction.
[0036] Select the largest eigenvalue The corresponding eigenvector This eigenvector is called the first principal component direction, which points to the direction in which the original data changes most dramatically, that is, the direction with the largest variance. Projecting onto the first principal component direction, thus generating a one-dimensional dynamics-dominated time series The calculation formula is: ; in, It's time of dimensional feature vector, is the first principal component eigenvector, Is a scalar value. By performing this projection operation on the feature vectors of all moments, a complete one-dimensional time series is obtained. This sequence serves as a comprehensive indicator, and its fluctuations in value condense the core dynamic pattern of the coordinated changes of multiple original acoustic features. This sequence will serve as the only input for the next step of phase space reconstruction.
[0037] S203, using the dynamics-dominant time series to perform phase space reconstruction to obtain a phase space trajectory representing the singer's dynamic state; The theoretical basis of this step is Takens embedding theorem, which states that for a deterministic dynamical system, the topological structure of the original system dynamical attractor can be restored by delaying the embedding of a single time series generated by the system. This allows us to recover the topological structure of the original system dynamical attractor from a one-dimensional observation signal. In this paper, a high-dimensional state space is reconstructed, and the trajectories in this space can reveal the intrinsic dynamics of the system.
[0038] In a specific embodiment, before performing phase space reconstruction, two key parameters need to be determined: embedding dimension and time delay The choice of these two parameters directly affects whether the reconstructed phase space can accurately reflect the characteristics of the original dynamic system.
[0039] Time Delay Determination of: Time delay Determines the time interval between the components used to construct the state vector. If it is too small, the components are highly correlated, and the reconstructed trajectory will be concentrated near the main diagonal and cannot be expanded; If the value is too large, the components may lose their correlation, which may lead to the introduction of noise in the trajectory. In a specific embodiment, the average mutual information method is used to determine the optimal .calculate and The mutual information between The function of the change, and select the first local minimum point of the function value as the optimal time delay.
[0040] Embedding Dimension Determination of: Embedding Dimension must be large enough to fully expand the attractor of the dynamical system. According to the theorem, if the attractor dimension of the original system is ,but In a specific embodiment, the pseudo-nearest neighbor method is used to determine the minimum effective embedding dimension. This method increases the embedding dimension by , and calculated in The nearest neighbor points in the dimensional space are Whether the dimension is sufficient is determined by whether there are still neighbors in the dimensional space. When the proportion of pseudo neighbor points drops to zero or a sufficiently small threshold, the corresponding The value is what you want.
[0041] Determined embedding dimension and time delay After that, the dynamics-dominated time series can be Perform embedding operations. At any time ,one dimensional dynamical state vector Constructed by the following formula: ; in, is a point in the reconstructed phase space whose coordinates are determined by the current moment and the past Equally spaced moments This vector captures the system at time The state of the local historical information and its evolution over time.
[0042] Over time The continuous advancement of the calculated dynamic state vector exist A continuous trajectory is depicted in the dimensional phase space. The geometry, density distribution, and evolution path of this phase space trajectory intuitively reflect the dynamic process of the singer's vocal system transitioning from one state to another. This trajectory is the direct target for subsequent benchmark model construction and real-time analysis.
[0043] S204. Based on the preset sample singing data, construct a two-stage dynamics benchmark model in the phase space. The two-stage dynamics benchmark model includes an ideal precursor trajectory cluster and an ideal take-off point for characterizing the vocal preparation stage, and a benchmark artistic manifold for characterizing the vocal performance stage. This step is an offline calculation process, the purpose of which is to generate a standardized dynamic reference system for comparison in subsequent real-time analysis. The benchmark model contains dynamic standards for the vocal preparation phase and the vocal performance phase.
[0044] In a specific embodiment, this step first requires preparing sample performance data, i.e., high-fidelity recordings of multiple opera singers performing the same aria. For each sample recording, steps S201 to S203 are first executed to convert it into its own phase space trajectory. Subsequently, a high-precision note onset detection algorithm is applied to the audio signal of each sample recording to obtain the onset timestamp of each note in the aria. ,in The serial number of the note.
[0045] Based on the above timestamps and phase space trajectories, a baseline model of the utterance preparation phase is constructed. This model includes an ideal starting point and an ideal precursor trajectory cluster. The ideal starting point Defined as the template at the beginning of the note To increase the robustness of the model, the final It can be the geometric center or mean of the coordinates of all templates at this point. Its corresponding ideal precursor trajectory cluster , all templates are preceded by a preset time window before the note start point The trajectory cluster is geometrically formed by the phase space trajectory segments within a period of 100 milliseconds. A manifold channel with a certain volume at the end point, which contains the dynamic shape of all the correct preparation actions performed by the template to play the note.
[0046] Next, we build a baseline model for the vocal performance stage, the baseline art manifold. First, all phase space state points in all sample singing data that are during the note duration (i.e., the time stamp is between any two consecutive note starting points) are extracted. These large number of state points from different samples and different notes together constitute a high-dimensional phase space point cloud set. This point cloud represents the set of all correct, artistically expressive vocalization states.
[0047] Since the inherent degrees of freedom of the high-dimensional point cloud are much lower than the dimensionality of the space in which it resides, a manifold learning algorithm is used to discover its inherent low-dimensional geometric structure. In a specific embodiment, the Isomap algorithm is used. The execution process of the algorithm includes: Point cloud collection For each point in , find its Euclidean distance nearest neighbor points and construct an adjacency graph where the edge weights are the Euclidean distances between points.
[0048] On this adjacency graph, the shortest path distance between any two points is calculated. This distance is used as an approximation of the geodesic distance between the two points, which reflects the true distance between the points on the manifold rather than the straight-line distance through high-dimensional space.
[0049] The calculated geodesic distance matrix is used as the input of the multidimensional scaling (MDS) algorithm. The MDS algorithm achieves distance-preserving embedding of the original point cloud by finding a low-dimensional space so that the Euclidean distance of the points in the space is as consistent as possible with the input geodesic distance. The final output low-dimensional embedding result is the benchmark art manifold .
[0050] After construction is completed, the two-stage dynamic benchmark model including the ideal starting points of all notes, the ideal precursor trajectory clusters and the benchmark artistic manifold is structured and stored in the benchmark model storage module 30 for real-time calling and comparison in step S205.
[0051] S205: Generate targeted singing guidance information based on comparative analysis of the phase space trajectory generated by the audio signal to be analyzed and the two-stage dynamics benchmark model.
[0052] This step is a real-time online processing process, which calls different parts of the benchmark model and performs corresponding analysis and guidance generation algorithms according to the different stages of the singing.
[0053] Voice preparation stage: Calculate and predict the starting point and the output of the PD controller. The entire process only involves basic vector addition, subtraction, multiplication and division operations, and the amount of calculation is extremely low.
[0054] During the vocalization phase, the deviation between the user's state point and the benchmark art manifold is calculated. The bottleneck lies in finding the nearest point on the manifold. By pre-establishing an efficient index structure during offline benchmark model construction, the online query process for the nearest point can be greatly accelerated, with computation time far below the perceptual latency of the human ear.
[0055] The system can generate phase space trajectories in real time and point by point as the user sings, and instantly compare them with the benchmark model, thereby providing continuous and uninterrupted guidance information, fully meeting the requirements of real-time feedback.
[0056] Phase 1: Feedforward Guidance in Vocal Preparation In a specific embodiment, when the system determines that the current When the preset sound preparation time window of the first note is within the preset sound preparation time window, the guidance logic of this stage will be triggered. The ideal starting point for each note and the ideal precursor trajectory cluster .
[0057] The core of this stage is prediction. The system continuously tracks the phase space trajectory points currently generated by the user. , and based on its recent motion state, predict the predicted starting point it will reach at the start of the note In one embodiment, the prediction is achieved by linear extrapolation: First, the current velocity vector of the user trajectory is calculated , and then based on the remaining preparation time Extrapolate to get the predicted take-off point .
[0058] After obtaining the predicted take-off point, the system calculates the error vector between it and the ideal take-off point. Based on this error, a feedforward guidance vector is generated. , which in one embodiment is generated by a proportional-derivative controller: ; Among them, the proportional term Provides a correction force proportional to the prediction error to pull the predicted trajectory toward the target; the differential term Damping is provided based on the rate of change of the error to suppress oscillations during the correction process and make the preparatory action smoother.
[0059] Phase 2: Feedback guidance during the vocal performance phase In a specific embodiment, when the system determines that the current state is in the duration of a note, the guidance logic of this stage will be triggered. The system reads the reference art manifold from the reference model storage module 30 .
[0060] For the user's current phase space state point , the system first calculates its This calculation is done by finding a distance on the manifold The nearest point is The orthogonal projection point on the manifold is denoted by The shortest distance This is the deviation measure.
[0061] Based on the relationship between the user's state point and the projection point, feedback guidance information is generated. This information can include two aspects: State Correction Guide: Generates a correction vector The vector points from the user's current point to the projection point on the manifold. Its direction and magnitude directly indicate the direction and magnitude in which the user needs to adjust his or her vocal state in order to return his or her dynamic state to the standard paradigm.
[0062] Tempo and Dynamic Guidance: This system generates guidance on tempo and dynamic strength by comparing the tangential velocity of the user's trajectory with the average velocity of the benchmark art manifold in that area. If the tangential velocity component of the user's trajectory exceeds the average velocity of the model, a deceleration instruction is generated; vice versa.
[0063] Ultimately, whether it is the feedforward guidance vector or the feedback guidance information, the system will reversely map it from the abstract phase space vector back to the specific acoustic feature dimension, and convert it into a visual graph (such as displaying the deviation of the trajectory from the target on the screen) or text instructions, and present it to the user.
[0064] Refer to the attached Figure 2 The present invention provides a multi-feature fusion opera vocal modeling system. This system can be deployed on computing devices, including but not limited to servers, personal computers, or dedicated embedded devices. The system may include a signal processing module 10, a dynamics modeling module 20, a baseline model storage module 30, and an analysis and guidance module 40.
[0065] The signal processing module 10 is configured to obtain an audio signal to be analyzed and process the signal to extract a multidimensional feature vector. In a specific embodiment, the signal processing module 10 first receives an original audio signal and applies a signal purification algorithm to remove the non-human voice part therein, thereby obtaining a pure human voice signal. Subsequently, the module divides the pure human voice signal into a series of overlapping short-time frames and extracts an N-dimensional acoustic feature vector for each frame. The vector includes features such as fundamental frequency, harmonic-to-noise ratio, spectral centroid, and Mel-frequency cepstral coefficients. The final output of the module is a series of continuous multidimensional feature vectors representing the singing process.
[0066] The dynamic modeling module 20 is connected to the signal processing module 10. This module receives the multi-dimensional feature vector sequence output by the signal processing module 10 and converts it into a dynamic trajectory in the phase space. In a specific embodiment, the dynamic modeling module 20 first performs principal component analysis to project the multi-dimensional feature vector sequence onto the first principal component direction to generate a one-dimensional dynamic dominant time series. Then, the module reconstructs the phase space based on the dynamic dominant time series. This process constructs a dimensional dynamical state vector : ; in, For the moment The dynamical state vector of For the moment The dynamics of the dominant time series, is the embedding dimension, The dynamics modeling module 20 continuously outputs a phase space trajectory consisting of continuous dynamics state vectors.
[0067] The benchmark model storage module 30 is a data storage unit configured to store a pre-calculated and constructed two-stage dynamic benchmark model. This benchmark model is generated offline based on multiple sets of model performance data. Its contents include: multiple ideal take-off points corresponding to specific notes, a cluster of ideal precursor trajectories associated with each ideal take-off point, and a benchmark artistic manifold representing the vocal performance phase of the entire aria. This module provides the stored benchmark model data to the analysis and guidance module 40.
[0068] The analysis and guidance module 40 is connected to the dynamics modeling module 20 and the benchmark model storage module 30. This module is the core analysis and decision-making unit of the system. It receives the real-time user phase space trajectory generated by the dynamics modeling module 20 and retrieves the corresponding benchmark model data from the benchmark model storage module 30 for comparative analysis. In a specific embodiment, this module performs two different operations depending on the stage of the performance: In the utterance preparation phase, the module predicts the user's upcoming starting point based on the user's preceding trajectory. , and generates a feedforward guidance vector : ; in, is the first The ideal starting point for each note, is the proportional gain coefficient, is the differential gain coefficient.
[0069] During the vocalization phase, the module calculates the deviation between the user's current state point and the reference art manifold and generates a correction vector for feedback. The module ultimately outputs this guidance information to a user interface or other output device.
[0070] During the operation of the entire system, each module works collaboratively. The signal processing module 10 is responsible for initial data acquisition and characterization. The dynamic modeling module 20 converts high-dimensional feature data into a low-dimensional representation of the dynamic state. The benchmark model storage module 30 provides the reference standard required for analysis. The analysis and guidance module 40 performs the core comparative analysis and guidance information generation tasks, forming a complete data processing and analysis process.
[0071] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A multi-feature fusion opera aria modeling method, characterized by: The following steps are involved: Acquire the audio signal to be analyzed and extract a multidimensional feature vector reflecting the singing state from it; generating a one-dimensional dynamics-dominated time series based on the multi-dimensional eigenvector; Reconstructing the phase space using the dynamics-dominant time series to obtain a phase space trajectory representing the singer's dynamic state; Based on preset model singing data, a two-stage dynamic benchmark model is constructed in phase space. The two-stage dynamic benchmark model includes an ideal precursor trajectory cluster and an ideal take-off point for characterizing the vocal preparation stage, and a benchmark artistic manifold for characterizing the vocal performance stage. Targeted singing guidance information is generated based on a comparative analysis of the phase space trajectory generated by the audio signal to be analyzed and the two-stage dynamics benchmark model.
2. The opera aria modeling method based on multi-feature fusion according to claim 1 is characterized in that: The step of generating a one-dimensional dynamics-dominated time series based on the multi-dimensional feature vector comprises: Constructing a data matrix from the continuous multi-dimensional feature vectors; Performing principal component analysis on the data matrix to obtain a first principal component direction; The multidimensional feature vector sequence is projected onto the first principal component direction to obtain the dynamics-dominated time series.
3. The opera aria modeling method based on multi-feature fusion according to claim 1 is characterized in that: The step of reconstructing the phase space using the dynamics-dominant time series comprises: Determining the embedding dimension and time delay ; At the moment , by the dynamics dominating the time series Perform time delay embedding to construct dynamic state vector ,in: ; in, For the moment The dynamical state vector of For the moment The dynamics of the dominant time series, is the embedding dimension, For time delay; The phase space trajectory is composed of the continuous dynamical state vector constitute.
4. The opera aria modeling method based on multi-feature fusion according to claim 1, characterized in that: The steps of constructing the two-stage kinetic benchmark model include: Perform note starting point detection on the sample singing data to obtain a note starting point timestamp; Marking the phase space state point corresponding to the timestamp of the note starting point as the ideal starting point; Constructing the phase space trajectory set within a preset time window before the timestamp of the note start point into the ideal precursor trajectory cluster; The phase space trajectory point cloud during the duration of the note is processed through a manifold learning algorithm to construct the benchmark art manifold.
5. The opera aria modeling method based on multi-feature fusion according to claim 1 is characterized in that: The step of generating targeted singing guidance information includes: In the utterance preparation phase, the phase space trajectory generated by the user is tracked as the user's precursor trajectory, and the predicted starting point to be reached by the user is predicted based on the user's precursor trajectory; The difference between the predicted starting point and the ideal starting point is calculated to generate feedforward guidance information for guiding the user's front-drive trajectory to approach the ideal starting point.
6. The opera aria modeling method based on multi-feature fusion according to claim 1 is characterized in that: The feedforward guidance information is a feedforward guidance vector , which is calculated as: ; in, For the moment The feedforward guidance vector, For the The ideal starting point for each note, For the moment For the first The predicted starting point of each note, is the proportional gain coefficient, is the differential gain coefficient, is a standard mathematical operator.
7. The opera aria modeling method based on multi-feature fusion according to claim 1 is characterized in that: The step of generating targeted singing guidance information also includes: During the vocalization stage, the user's current phase space state point is obtained; Calculating the deviation of the phase space state point from the reference art manifold and finding the projection point on the reference art manifold; Based on the relationship between the phase space state point and the projection point, feedback guidance information is generated for guiding the phase space state point to return to the reference art manifold.
8. The opera aria modeling method of multi-feature fusion according to claim 1 is characterized in that: The feedback guidance information includes a correction vector for pulling the deviated phase space state point back to the reference art manifold, and the correction guides the user to adjust the vocal state so that its trajectory in the phase space is close to the reference art manifold.
9. The opera aria modeling method of multi-feature fusion according to claim 1 is characterized in that: The manifold learning algorithm is an Isomap algorithm, which discovers and constructs the benchmark art manifold by calculating the geodesic distance between points in the phase space trajectory point cloud and applying multidimensional scaling analysis.
10. A multi-feature fusion opera aria modeling system, according to the multi-feature fusion opera aria modeling method according to any one of claims 1-9, characterized in that: include: A signal processing module is used to obtain the audio signal to be analyzed and extract a multidimensional feature vector reflecting the singing state; A dynamic modeling module, configured to generate a one-dimensional dynamic dominant time series based on the multidimensional feature vector, and to perform phase space reconstruction using the dynamic dominant time series to obtain a phase space trajectory representing the dynamic state of the singer; A benchmark model storage module is used to store a preset two-stage dynamic benchmark model constructed in phase space, wherein the two-stage dynamic benchmark model includes an ideal precursor trajectory cluster and an ideal take-off point for characterizing the vocal preparation stage, and a benchmark art manifold for characterizing the vocal performance stage; The analysis and guidance module is used to generate targeted singing guidance information based on the comparative analysis of the phase space trajectory generated by the dynamic modeling module and the two-stage dynamic benchmark model stored in the benchmark model storage module.
Citation Information
Cited By
Belt weigher remote automatic metering system and method
CN121409382A