A deep learning-based vocalization skill recommendation method

By simulating the biomechanics of vocal organs through a brain-like acoustic perception layer and a digital twin model, and combining adversarial aesthetic assessment and a cross-modal metaphor generation system, the system solves the problem of the singularity of existing vocal technique recommendation systems, achieves accurate and personalized vocal technique recommendations, and improves the user's training effect.

CN122224142APending Publication Date: 2026-06-16GUIZHOU RADIO & TV UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUIZHOU RADIO & TV UNIV
Filing Date
2026-03-13
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing vocal technique recommendation systems lack in-depth integration of users' physiological conditions, artistic goals, and personalized perceptual needs, resulting in recommendations that are either too scientific or too artistic, failing to meet the dual needs of professional vocal training and general learning.

Method used

By using brain-like acoustic perception layer hierarchical coding and digital twin model to simulate the biomechanics of vocal organs, combined with adversarial aesthetic assessment and cross-modal metaphor generation system, multi-dimensional data is integrated to recommend vocal techniques and dynamically adjust modal weights to adapt to user feedback.

Benefits of technology

It achieves precise matching of vocal technique recommendations, improves the accuracy and personalization of recommendations, provides a combination of scientific and artistic techniques, and enhances the user training experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122224142A_ABST
    Figure CN122224142A_ABST
Patent Text Reader

Abstract

The application provides a vocalization skill recommendation method based on deep learning, which comprises the following steps: layering coding of original sound waves of a user to obtain an acoustic feature vector; construction of a digital twin model of a vocal organ to simulate vocalization changes and effects; establishment of an adversarial aesthetic evaluation system to generate a vocalization skill combination; construction of a cross-modal metaphor generation system to generate personalized training guidance and convert it into a concrete metaphor description and supporting visual materials; and fusion of the acoustic feature vector, biomechanical changes and other materials by using a multi-modal fusion algorithm to generate a vocalization skill recommendation result adapted to the user. The application integrates multi-dimensional data to construct a three-dimensional feature space, avoids the limitations of single-dimensional analysis, dynamically adjusts the weights of each mode by using a gating unit, optimizes the recommendation strategy in combination with user feedback, unifies professional rules and public preferences through the adversarial aesthetic evaluation system, generates a skill combination with scientific and artistic qualities, and improves the accuracy and personalization level of vocalization skill recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vocal technique recommendation technology, and in particular to a vocal technique recommendation method based on deep learning. Background Technology

[0002] With the integration of vocal education and intelligent technology, vocal technique recommendation has become a key technology for improving learners' efficiency. Currently, vocal technique recommendations mainly rely on two models: one is the subjective guidance of teachers based on their personal experience in traditional vocal teaching, and the other is a recommendation system based on rules or simple statistical models. Traditional methods are limited by the differences in individual teacher levels, and the recommendation results are highly subjective and difficult to standardize; while existing intelligent recommendation systems mostly focus on the single-dimensional analysis of acoustic parameters (such as pitch and intensity), or achieve technique recommendations through shallow models such as collaborative filtering, lacking a deep integration of users' physiological conditions, artistic goals, and personalized perceptual needs.

[0003] Existing technologies rely too heavily on expert subjective judgment or superficial acoustic characteristics for recommendations, lacking multimodal fusion analysis of vocal organ biomechanics, user physiological data, and artistic expression. Secondly, existing systems fail to establish a unified evaluation framework that integrates professional acoustic principles with public listening preferences, resulting in recommendations that either prioritize scientific accuracy at the expense of user experience or excessively pursue artistry without professional support. Finally, traditional methods lack dynamic adaptation mechanisms, failing to adjust recommendation strategies based on real-time user feedback, and personalized training methods such as metaphorical guidance have not yet been deeply integrated with intelligent recommendation systems. These shortcomings make it difficult for existing technologies to meet the dual needs of professional vocal training and general learning in terms of accuracy, personalization, and user experience. Summary of the Invention

[0004] This invention aims to at least address the problem of low recommendation accuracy in existing technologies, and innovatively proposes a deep learning-based method for recommending vocal techniques.

[0005] To achieve the above-mentioned objectives of this invention, this invention provides a deep learning-based method for recommending vocal techniques, the method comprising: S1. The original sound waves emitted by the user are encoded in layers through the brain-like acoustic perception layer to obtain acoustic feature vectors. S2. Construct a digital twin model of the user's vocal organs based on the user's physiological data, and generate personalized 3D biomechanical models of the vocal cords, larynx, and oral cavity through finite element analysis and fluid dynamics simulation. Simulate the biomechanical changes and acoustic effects during vocalization through a real-time physics engine. S3. Establish an adversarial aesthetic evaluation system, which includes a gold standard dataset and a public preference comparison learning framework. Through the multi-task public preference comparison learning framework, the acoustic rules of professional judges and the public's listening preferences are mapped to the same semantic space to generate a combination of vocal techniques that are both scientific and artistic. S4. Construct a cross-modal metaphor generation system based on vocal technique tags, establish a knowledge graph of vocal techniques and figurative metaphors based on natural language processing methods, generate personalized training guidance based on the knowledge graph through a diffusion model, and transform the abstract personalized training guidance into perceptible figurative metaphor descriptions and supporting visual materials. S5. The acoustic feature vectors, biomechanical changes, acoustic effects, vocal technique combinations, metaphorical descriptions, and accompanying visual materials are fused using a multimodal fusion algorithm to generate vocal technique recommendations that are adapted to the user's physiological conditions and artistic goals.

[0006] The beneficial effects of this invention are as follows: First, it integrates multi-dimensional data such as brain-like acoustic features, biomechanical parameters, acoustic effects, and personalized metaphorical descriptions to construct a feature space covering physiological, acoustic, and artistic dimensions, avoiding the limitations of single-dimensional analysis. Second, it uses gating units to dynamically adjust the weights of each modality (e.g., strengthening biomechanical features for users with sensitive vocal cords, and increasing the weight of acoustic effects for popular singers), and optimizes the recommendation strategy in real time based on user feedback. Finally, it unifies professional acoustic principles with public preferences through adversarial aesthetic evaluation, generating skill combinations that are both scientific and artistic. Compared to traditional recommendation methods that rely on subjective experience or shallow models, this invention achieves precise adaptation throughout the entire process from data collection and feature fusion to result generation, improving the accuracy and personalization of vocal technique recommendations.

[0007] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0008] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart of a deep learning-based method for recommending vocal techniques according to the present invention. Detailed Implementation

[0009] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0010] like Figure 1 As shown, a deep learning-based method for recommending vocal techniques includes: S1. The original sound waves emitted by the user are encoded in layers through the brain-like acoustic perception layer to obtain acoustic feature vectors. S2. Construct a digital twin model of the user's vocal organs based on the user's physiological data, and generate personalized 3D biomechanical models of the vocal cords, larynx, and oral cavity through finite element analysis and fluid dynamics simulation. Simulate the biomechanical changes and acoustic effects during vocalization through a real-time physics engine. S3. Establish an adversarial aesthetic evaluation system, which includes a gold standard dataset and a public preference comparison learning framework. Through the multi-task public preference comparison learning framework, the acoustic rules of professional judges and the public's listening preferences are mapped to the same semantic space to generate a combination of vocal techniques that are both scientific and artistic. S4. Construct a cross-modal metaphor generation system based on vocal technique tags, establish a knowledge graph of vocal techniques and figurative metaphors based on natural language processing methods, generate personalized training guidance based on the knowledge graph through a diffusion model, and transform the abstract personalized training guidance into perceptible figurative metaphor descriptions and supporting visual materials. S5. The acoustic feature vectors, biomechanical changes, acoustic effects, vocal technique combinations, metaphorical descriptions, and accompanying visual materials are fused using a multimodal fusion algorithm to generate vocal technique recommendations that are adapted to the user's physiological conditions and artistic goals.

[0011] In step S5, it is necessary to explain in detail that the multimodal fusion algorithm first preprocesses the acoustic feature vectors, biomechanical changes, and acoustic effects to ensure the consistency of each modality's data in time and space. Subsequently, the algorithm uses an attention mechanism to dynamically adjust the weights of different modalities to reflect the importance of the user's specific physiological conditions and artistic goals. For example, for users with sensitive vocal cords, the weight of biomechanical changes will be increased accordingly. Furthermore, the modal fusion algorithm in this embodiment also considers the semantic association between vocal technique combinations and metaphorical descriptions, tightly combining these two through semantic embedding technology. Finally, through optimization of the multimodal fusion loss function, the modal fusion algorithm generates a vocal technique recommendation result that conforms to both the user's physiological conditions and their artistic goals. This result not only includes scientific vocal techniques but also incorporates artistic expression.

[0012] The technical principle of the deep learning-based voice technique recommendation method in this embodiment is as follows: First, in step S1, the brain-like acoustic perception layer performs fine-grained layered encoding on the raw sound waves emitted by the user. This process simulates the complex sound processing mechanism of the human brain, thereby capturing more nuanced and comprehensive acoustic features. These feature vectors not only contain basic information such as the frequency and amplitude of the sound, but also contain advanced features such as the timbre and tone quality of the sound.

[0013] Next, in step S2, a highly personalized 3D biomechanical model is generated by constructing a digital twin model of the user's vocal organs and combining finite element analysis and fluid dynamics simulation techniques. This model can simulate the biomechanical changes during vocalization in real time, including the vibration patterns of the vocal cords, and the shape changes of the larynx and oral cavity, thus accurately reflecting the physiological mechanisms of the user's vocalization. This process not only improves the accuracy of vocal technique recommendations but also provides users with more scientific and personalized training guidance.

[0014] In step S3, the established adversarial aesthetic evaluation system is a crucial step in combining the acoustic principles of professional reviewers with the listening preferences of the general public. Through a multi-task public preference comparison learning framework, this system maps both to the same semantic space, thereby generating a combination of vocal techniques that is both scientifically sound and artistically expressive. This process not only ensures the scientific validity of the vocal techniques but also fully considers the listening preferences of the general public, making the recommended vocal techniques more closely aligned with users' actual needs.

[0015] Furthermore, in step S4, through natural language processing methods and knowledge graph technology, the system can establish a correspondence between vocal techniques and concrete metaphors, and generate personalized training guidance text through a diffusion model. Subsequently, through semantic analysis and sentiment recognition technology, visual materials matching the vocal techniques and concrete metaphors are selected from a pre-set visual material library, ultimately generating a description of the concrete metaphor and accompanying visual materials. This process not only improves the user's understanding and mastery of vocal techniques but also enhances the fun and interactivity of the training.

[0016] In summary, this invention's deep learning-based vocal technique recommendation method achieves precise adaptation throughout the entire process, from data collection and feature fusion to result generation, through innovative means such as integrating multi-dimensional data, utilizing deep learning technology for feature extraction and accurate understanding, and constructing a cross-modal metaphor generation system. This method not only improves the accuracy and personalization of vocal technique recommendations but also provides users with a more scientific and engaging training experience.

[0017] As an optional embodiment of the present invention, the brain-like acoustic perception layer in step S1 may include a spiking neural network, which extracts the spatiotemporal pulse pattern, time-related features and abstract acoustic features of sound waves through a primary layer, an intermediate layer and a high-level layer, respectively.

[0018] It should be noted that the spiking neural network in this embodiment employs an adaptive learning rate adjustment strategy to optimize the training process of network parameters. Specifically, this strategy dynamically adjusts the learning rate based on the loss function value during training. When the loss function value decreases slowly, the learning rate is appropriately reduced to avoid overfitting; conversely, when the loss function value decreases rapidly, the learning rate is appropriately increased to accelerate convergence. Furthermore, this embodiment also introduces a regularization term to further prevent overfitting and improve the model's generalization ability. By employing these techniques, the spiking neural network in this embodiment can extract acoustic wave features more accurately.

[0019] The spiking neural network in this embodiment includes a primary layer, an intermediate layer, and a high-level layer. The primary layer is mainly responsible for extracting the spatiotemporal pulse patterns of sound waves. These patterns reflect the basic physical characteristics of sound waves, such as frequency, amplitude, and waveform. The intermediate layer further processes these spatiotemporal pulse patterns to extract time-related features. These features reveal the dynamic changes of sound waves over time, such as pitch fluctuations and rhythmic variations. Finally, the high-level layer further abstracts the time-related features extracted by the intermediate layer to generate abstract acoustic features. These features include advanced attributes such as sound quality and timbre, which can more comprehensively reflect the aesthetic characteristics of sound.

[0020] As an optional embodiment of the present invention, the expression of the primary layer may be: in, Indicates the first A neuron in time membrane potential, This represents the time step in a discrete-time simulation. This represents the membrane potential decay time constant. This represents the input current after sound wave conversion. Indicates the first Synaptic weights of individual neurons Represents neurons Pulse firing status, This represents the neuron firing threshold, where 1 indicates firing and 0 indicates resting. The expression for the intermediate layer is: in, Represents presynaptic neurons to postsynaptic neurons The change in connection weights, This represents the learning rate within the positive time window. Indicates the pulse timing difference. This represents the decay constant of the positive time window. This represents the decay constant for the negative time window. Represents time-related feature vectors. This indicates a pooling operation. Represents postsynaptic neurons The pulse state; The expression for the higher-level layer is: in, This represents the abstract acoustic feature vector output by the higher-level layer. Represents a non-linear activation function. Indicates the first Synaptic weights of individual neurons Indicates the first A neuron in the time window Weighted pulse count within, Indicates the first Bias terms for each neuron, Indicates the first A neuron in time The pulse state at that time, This represents the pulse count decay time constant.

[0021] As an optional embodiment of the present invention, the simulation of biomechanical changes and acoustic effects during sound production using a real-time physics engine in step S2 may include: S201. Preprocess the user's physiological data; In step S201, it should be noted that preprocessing includes noise removal, data smoothing, and standardization. This step helps to eliminate errors and interference that may occur during the physiological data acquisition process.

[0022] S202. Based on the preprocessed physiological data of the user, establish a finite element model of vocal cord biomechanics, a finite element model of laryngeal fluid, and a finite element model of oral cavity tuning; In step S202, it should be noted that the steps for establishing the vocal cord biomechanical finite element model are as follows: First, a geometric model of the vocal cords is constructed using finite element analysis software based on the user's vocal cord morphology, muscle distribution, and vocal cord vibration characteristics. Then, material properties, such as elastic modulus and density, are assigned to each element in the geometric model to simulate the biomechanical properties of the vocal cords. Next, boundary conditions and loads, such as tension on both sides of the vocal cords and the impact force of airflow on the vocal cords, are set to simulate the biomechanical changes during phonation. The finite element solver is then used to calculate the biomechanical parameters such as stress, strain, and displacement during vocal cord vibration.

[0023] Similarly, the construction process for the finite element model of laryngeal fluid and the finite element model of oral cavity tuning is similar. The laryngeal fluid finite element model needs to simulate the flow of air within the laryngeal cavity, including fluid dynamic parameters such as velocity, pressure, and vorticity. The oral cavity tuning finite element model, on the other hand, needs to simulate the impact of changes in oral cavity shape on sound, including parameters such as the shape, size, and location of the resonance cavities. These models together constitute a digital twin model of the user's vocal organs, capable of simulating the biomechanical changes and acoustic effects during vocalization in real time.

[0024] S203. The vocal cord biomechanical finite element model, the laryngeal cavity fluid finite element model and the oral cavity tuning finite element model are bidirectionally coupled to obtain a digital twin model of the user's vocal organs. In step S203, it should be noted that the specific steps for establishing a digital twin model of the user's vocal organs in this embodiment are as follows: First, the geometric coordinates and physical parameters of the vocal cord biomechanical finite element model, the laryngeal cavity fluid finite element model, and the oral cavity tuning finite element model are unified to ensure seamless connection between the models.

[0025] Subsequently, using an existing multiphysics coupling algorithm, the three models were bidirectionally coupled in time and space. In this process, the vocal cord biomechanics finite element model provides vibration information of the vocal cords, the laryngeal fluid finite element model simulates the impact of airflow on the vocal cords and the fluid dynamics changes within the laryngeal cavity, while the oral cavity tuning finite element model adjusts the resonance effect of the sound according to changes in the shape of the oral cavity. Through calculations using the multiphysics coupling algorithm, a highly personalized digital twin model of the user's vocal organs was finally obtained. This model can accurately simulate the biomechanical changes and acoustic effects during the user's vocalization in real time.

[0026] S204. Import the digital twin model into the real-time physics engine and optimize the model parameters of the digital twin model using the Bayesian optimization algorithm. In step S204, it should be noted that the Bayesian optimization algorithm iteratively adjusts the model parameters to minimize the difference between the simulation results and the actual sound output. Specifically, the algorithm first generates a set of simulation results based on the current parameter settings and calculates the error between these results and the preset target. Then, the algorithm uses Bayes' theorem to update the posterior distribution of the parameters to determine the settings for the next set of parameters. This process is repeated until a preset convergence condition or an upper limit on the number of iterations is reached. By employing the Bayesian optimization algorithm, this embodiment can efficiently optimize the parameters of the digital twin model, thereby improving the accuracy and real-time performance of the simulation.

[0027] The Bayesian optimization algorithm described above is an existing Bayesian optimization algorithm used in this embodiment.

[0028] S205. Adjust the vocal cord biomechanical finite element model in the digital twin model using a non-rigid registration algorithm; optimize the laryngeal cavity cross-sectional area change curve using airflow pressure data from the user's physiological data, and define a personalized relationship between subglottic pressure and glottal impedance; optimize the tongue trajectory and lip shape parameters in the oral cavity tuning finite element model based on vocalization task data using a genetic algorithm. In step S205, it is important to explain in detail that the application of the non-rigid registration algorithm aims to further improve the matching degree between the digital twin model and the user's actual vocalization. This algorithm can take into account the dynamic deformation of the vocal cords during the vocalization process, thereby adjusting the geometric shape and physical parameters of the vocal cord biomechanical finite element model to make it closer to the user's real physiological state.

[0029] The optimization of the laryngeal cavity cross-sectional area change curve is based on airflow pressure data from the user's physiological data. By analyzing the change pattern of airflow pressure over time in detail, the dynamic adjustment process of the laryngeal cavity during phonation can be simulated more accurately, thereby optimizing the cross-sectional area change curve and improving the accuracy of the simulation.

[0030] Furthermore, this embodiment defines a personalized relationship between subglottic pressure and glottal impedance. This relationship is established based on in-depth analysis of the user's physiological data and reflects the dynamic balance between subglottic pressure and glottal impedance during vocalization.

[0031] Furthermore, for the finite element model of oral vocalization, this embodiment utilizes a genetic algorithm to optimize the tongue trajectory and lip shape parameters. As a global optimization search method, the genetic algorithm can automatically acquire and accumulate knowledge about the search space and adaptively control the search process to obtain the optimal solution. By introducing a genetic algorithm, this embodiment can efficiently explore the optimal combination of tongue trajectory and lip shape parameters, thereby improving the accuracy and realism of the digital twin model in simulating user vocalization.

[0032] The specific steps for optimizing tongue trajectory and lip shape parameters using a genetic algorithm are as follows: First, based on the pronunciation task data, an objective function is defined to evaluate the difference between the simulated and actual vocalization effects under the current combination of tongue position trajectory and lip shape parameters. Next, the population is initialized by generating a set of random combinations of tongue position trajectory and lip shape parameters as the initial solution for the genetic algorithm. Then, the fitness of each individual in the population is evaluated, i.e., its objective function value is calculated. Based on the fitness evaluation results, superior individuals are selected as parents, and new offspring individuals are generated through genetic operations such as crossover and mutation. This process iterates until a preset convergence condition or an upper limit on the number of iterations is reached. Finally, the genetic algorithm outputs a set of optimal combinations of tongue position trajectory and lip shape parameters, significantly improving the accuracy and realism of the digital twin model in simulating user vocalization.

[0033] S206. Based on the optimized model parameters of the digital twin model, the biomechanical changes and acoustic effects during sound production are simulated using finite element analysis and fluid dynamics simulation methods.

[0034] In step S206, it is necessary to explain in detail that finite element analysis and fluid dynamics simulation methods can precisely simulate various complex biomechanical and acoustic phenomena during the sound production process. Finite element analysis can accurately calculate biomechanical parameters such as stress, strain, and displacement during vocal cord vibration, revealing the dynamic deformation and mechanical properties of the vocal cords during sound production. Simultaneously, fluid dynamics simulation methods can simulate the flow of air in the larynx and oral cavity, including fluid dynamic parameters such as velocity, pressure, and vorticity, thereby providing a deeper understanding of the influence of airflow on sound production.

[0035] As an optional embodiment of the present invention, optionally, the expression for the non-rigid registration algorithm in step S205 is: in, This represents the optimal displacement field. Indicates the number of corresponding point pairs. Indicates the first The spatial location of each coordinate point Indicates the first The original positions of the coordinate points Represents the displacement field function. This represents the regularization weight coefficient. Indicates the number of nodes in the target model. This represents the second-order gradient norm of the displacement field. Indicates the first The original positions of the coordinate points This represents the weight coefficient of the constraint term. The first part of the vocal cord after deformation One morphological feature, This represents the first position of the vocal cords in the user's physiological data. Reference values ​​for morphological features.

[0036] As an optional embodiment of the present invention, optionally, generating a combination of vocal techniques that combines scientific and artistic expression in step S3 includes: S301. Based on professional acoustic principles and public listening preference data, construct the gold standard dataset and the public preference dataset; In step S301, it is necessary to explain in detail that the gold standard dataset mainly covers professional knowledge in the field of acoustics, including objective acoustic parameters such as sound frequency, amplitude, and timbre, as well as the correspondence between these parameters and vocal techniques. This data comes from experimental measurements and analyses by acoustic experts, ensuring the scientific rigor and accuracy of the dataset.

[0037] The public preference dataset, on the other hand, focuses on collecting and analyzing public perceptions of different vocal techniques. Through questionnaires, online assessments, and other methods, it gathers a large number of listeners' subjective evaluations of different sound samples, thereby constructing a dataset reflecting public preferences. This dataset demonstrates the popularity and acceptance of vocal techniques in practical applications.

[0038] S302. Utilize a pre-trained acoustic model as a shared encoder, and extract professional acoustic features and public perception features based on the gold standard dataset and the public preference dataset, and output scientific and artistic scores respectively through a dual-branch network. In step S302, it is important to explain in detail that the pre-trained acoustic model, as a shared encoder, possesses powerful feature extraction capabilities, enabling it to capture rich acoustic features from sound signals. In this embodiment, this model is used to simultaneously process sound samples from both the gold standard dataset and the public preference dataset, extracting information that includes both professional acoustic features and public perception features. Subsequently, this information is input into a dual-branch network. The dual-branch network consists of two parallel neural network branches, responsible for outputting scientific and artistic scores, respectively. The scientific scoring branch evaluates based on professional acoustic features, objectively reflecting the scientific validity and accuracy of vocal techniques. The artistic scoring branch, on the other hand, evaluates based on public perception features, subjectively reflecting the artistic expression and popularity of vocal techniques.

[0039] The steps for training the acoustic model in this embodiment are as follows: First, a large number of sound samples are collected, covering various vocal techniques and voice styles. Then, these sound samples are preprocessed, including denoising and standardization, to ensure data quality. Next, a deep learning model is constructed as the initial architecture of the acoustic model. This model may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), or combinations thereof. Then, unsupervised or supervised learning methods are used to pre-train the model to learn the basic features and patterns in the sound signal. During pre-training, features such as the spectrogram and Mel-frequency cepstral coefficients (MFCCs) of the sound samples can be used as input, and the labels or categories of the sound samples can be used as output. After pre-training, sound samples from the gold standard dataset and the popular preference dataset are input into the pre-trained acoustic model to extract professional acoustic features and popular perception features. These features are then input into a two-branch network for further processing and scoring. Through continuous iteration and optimization of the training process, an acoustic model that can accurately evaluate the scientific and artistic expression of vocal techniques is finally obtained.

[0040] S303. Based on the scientific and artistic scores, an adversarial training strategy is used to distinguish between the professional feature distribution and the general feature distribution using a discriminator, thereby prompting the shared encoder to generate unified professional and general features that are independent of modality. In step S303, it is necessary to explain in detail that the adversarial training strategy aims to further improve the accuracy and generalization ability of the acoustic model in evaluating vocal skills. Specifically, this strategy introduces a discriminator whose task is to distinguish whether the input features come from a professional acoustic feature distribution or a general perception feature distribution. At the same time, the shared encoder is required to generate features that can confuse the discriminator, that is, to strive to make the generated features difficult to distinguish between the two distributions.

[0041] To achieve this goal, during training, the shared encoder continuously attempts to generate features that conform to both professional acoustic principles and resonate with general public perception. Meanwhile, the discriminator continuously learns and updates its discriminative capabilities to more accurately identify the source of the features. This adversarial process drives the shared encoder to continuously optimize its feature generation capabilities, ultimately generating unified professional and general public features independent of modality. These features not only possess scientific accuracy but also reflect the artistic preferences of the general public.

[0042] By employing an adversarial training strategy, this embodiment effectively addresses the modal discrepancy between professional acoustic features and generalized perception features, improving the accuracy and generalization ability of the acoustic model in evaluating vocal techniques. This will help recommend vocal technique combinations that better suit users' personal preferences and vocal needs, enhancing their vocal experience and effectiveness.

[0043] S304. By comparing and learning, the semantic distance between the professional features and the general features is narrowed, and by combining dynamic weighting with gating units, a fusion feature that combines both evaluation criteria is obtained. In step S304, it is important to explain in detail that the application of contrastive learning aims to further narrow the semantic gap between professional features and general features, enabling them to be integrated at a higher level. Specifically, contrastive learning constructs positive and negative sample pairs, learning how to shorten the distance between similar samples while widening the distance between dissimilar samples. In this embodiment, positive sample pairs consist of voice samples with similar vocal techniques and vocal styles, while negative sample pairs consist of voice samples with significant differences. Through contrastive learning, the model can learn the inherent connection between professional and general features, thereby achieving their integration at the semantic level.

[0044] The dynamic weighting mechanism of the gating unit is used to flexibly adjust the weights of professional and popular features during the fusion process. This mechanism dynamically calculates the weight of each feature based on the specific characteristics of the current input sound sample, ensuring that the fused features simultaneously reflect scientific accuracy and artistic expression. Specifically, the gating unit generates a set of weight coefficients based on the dimension and importance of the input features. These coefficients are then used for weighted summation to obtain the final fused features. By introducing the dynamic weighting mechanism of the gating unit, this embodiment can more precisely control the fusion process, improving the accuracy and practicality of the fused features.

[0045] Through the steps described above, this embodiment successfully constructs a fusion feature that combines scientific rigor with artistic expression. These fusion features not only reflect the objective acoustic characteristics of vocal techniques but also embody the public's subjective preferences for vocal techniques, thus enabling the recommendation of technique combinations that better suit users' individual needs and vocal characteristics.

[0046] S305, Based on the fusion features, the decoder is used to generate a combination of vocal techniques.

[0047] In step S305, it is necessary to explain in detail that the decoder's task is to convert the fused features into specific combinations of vocal techniques. In this embodiment, the decoder uses a deep learning model, such as a recurrent neural network (RNN) or a Transformer. These models have powerful sequence generation capabilities and can gradually generate corresponding combinations of vocal techniques based on the input feature sequence.

[0048] The training process of the decoder is closely related to the training process of the acoustic model. During the training phase, the decoder receives features extracted and fused by the acoustic model as input and attempts to generate vocal technique combinations that match these features. To evaluate the accuracy and reasonableness of the generated technique combinations, an auxiliary loss function can be introduced, which calculates the difference between the generated technique combinations and preset standard technique combinations. Through continuous iteration and optimization of the training process, the decoder can gradually learn how to generate high-quality vocal technique combinations based on the fused features.

[0049] In practical applications, when users request personalized vocal technique recommendations, the system first collects their physiological and vocal task data and constructs a digital twin model. Subsequently, it optimizes the digital twin model using Bayesian optimization algorithms, non-rigid registration algorithms, and genetic algorithms to more accurately reflect the user's actual vocalization. After model optimization, the system simulates the biomechanical changes and acoustic effects during vocalization using finite element analysis and fluid dynamics simulation methods to further verify the model's accuracy.

[0050] Next, the system inputs the features of the optimized digital twin model into the trained acoustic model, extracting professional acoustic features and general perception features, and outputs scientific and artistic scores respectively through a dual-branch network. Then, an adversarial training strategy is used to prompt the shared encoder to generate unified professional and general features independent of modality, and a fused feature that combines both evaluation criteria is obtained through contrastive learning and dynamic weighting mechanism of gating units.

[0051] Finally, the decoder generates specific combinations of vocal techniques based on the fusion features and recommends them to the user. These combinations not only conform to professional knowledge in the field of acoustics but also reflect the public's subjective preferences for vocal techniques, thus helping users better master vocal techniques and improve their vocal performance and expressiveness.

[0052] As an optional embodiment of the present invention, the expression of the adversarial training strategy in step S303 is optionally: in, Represents the adversarial training loss function. This represents a sample from the gold standard dataset. Indicates the distribution of professional data. Represents the mathematical expectation. Indicates the discriminator, Indicates a shared encoder. This represents a sample of a mass preference dataset. Indicates the distribution of mass data; The expression for obtaining the fusion feature that combines both evaluation criteria in step S304 is: in, This represents the fused feature vector. This represents the Sigmoid activation function. Represents the weight matrix of the gated unit. Represents the professional acoustic feature vector. Represents the feature vector of public perception. Indicates the gate unit bias term. This represents element-wise multiplication. This represents the contrastive learning semantic alignment loss. Represents the cosine similarity function. Indicates the temperature coefficient. Indicates the number of negative samples. Indicates the first The public's perceived feature vector of a negative sample This represents the weighting coefficient for the comparative loss.

[0053] As an optional embodiment of the present invention, optionally, in step S4, transforming the abstract personalized training guidance into a perceptible, concrete metaphorical description and accompanying visual materials includes: S401. Using natural language processing methods, extract keywords of vocal techniques and their corresponding concrete metaphorical descriptions from vocal music teaching literature, and construct a database of the correspondence between vocal techniques and concrete metaphors. In step S401, it is necessary to explain in detail that, in order to construct a database of correspondences between vocal techniques and concrete metaphors, the system first extensively collects vocal teaching literature, covering multiple aspects such as vocal theory, teaching practice, and analysis of vocal techniques. Subsequently, natural language processing technology is used to perform in-depth analysis of these documents, extracting keywords related to vocal techniques. These keywords represent different vocal techniques or methods.

[0054] While extracting keywords, the system further analyzes the descriptions of each keyword in the literature, especially those that use concrete metaphors. Concrete metaphors are a way of describing abstract concepts by transforming them into concrete images or scenes, helping users understand the key points and essence of vocal techniques more intuitively. For example, describing the abstract technique of "opening the throat" as "relaxing the throat like yawning" makes it easier for users to master this technique.

[0055] After collecting sufficient keywords and concrete metaphorical descriptions, the system organizes this information into a relational database. In this database, each vocal technique keyword is associated with one or more concrete metaphorical descriptions, forming a rich and intuitive knowledge base. When a user needs personalized vocal technique guidance, the system can retrieve the most suitable concrete metaphorical description from the database based on the user's actual situation and needs, and present it to the user.

[0056] In this way, the present invention not only transforms abstract, personalized training guidance into perceptible, concrete metaphorical descriptions, but also provides users with accompanying visual materials. These visual materials can be images, animations, or videos related to the concrete metaphors, which can further enhance users' understanding and mastery of vocal techniques. For example, when presenting the concrete metaphor of "relaxing the throat like yawning," the system can display an image or animation of yawning to help users more intuitively feel the relaxed state of the throat.

[0057] In summary, by constructing a database of correspondences between vocal techniques and concrete metaphors, and presenting it using natural language processing technology and visual materials, this invention provides users with richer, more intuitive, and personalized guidance on vocal techniques. This will help users better understand and master vocal techniques, improving vocal quality and expressiveness.

[0058] S402. Based on the aforementioned correspondence database, a knowledge graph of vocal techniques and figurative metaphors is constructed using a knowledge graph construction method, wherein vocal techniques are entity nodes in the knowledge graph, and figurative metaphors are relation edges associated with the entity nodes of vocal techniques. In step S402, it is necessary to explain in detail that the construction of the knowledge graph aims to further explore and demonstrate the intrinsic connection between vocal techniques and concrete metaphors. Specifically, a knowledge graph is a graphical method of knowledge representation that visually displays the relationships and hierarchical structure between knowledge by connecting entities and relationships in the form of nodes and edges.

[0059] In this embodiment, vocal techniques are represented as entity nodes in a knowledge graph, with each node representing a unique vocal technique or method. Concrete metaphors are represented as relational edges associated with the vocal technique entity nodes; each edge connects one or more vocal technique nodes and describes the concrete metaphorical relationships between these techniques.

[0060] To construct this knowledge graph, the system first traverses the corresponding relational database to extract vocal technique keywords and concrete metaphorical descriptions. Then, using knowledge graph construction methods, such as deep learning models or graph databases, these keywords and descriptions are transformed into nodes and edges in the knowledge graph. During the construction process, the system considers the similarities and connections between vocal techniques, as well as the logical relationships and hierarchical structure between concrete metaphors, to ensure that the generated knowledge graph is both accurate and easy to understand.

[0061] Once the knowledge graph is constructed, the system can provide personalized vocal technique guidance based on the user's actual needs. For example, when a user wants to learn a specific vocal technique, the system can locate the corresponding node in the knowledge graph and display its associated concrete metaphorical description. Simultaneously, the system can dynamically adjust the presentation and content of the knowledge graph based on user feedback and learning progress to provide guidance that better suits the user's needs.

[0062] The construction of a knowledge graph also facilitates intelligent recommendation and correlation analysis of vocal techniques. By analyzing and mining the nodes and edges in the knowledge graph, the system can discover potential connections and patterns between different vocal techniques, thereby recommending skill combinations that better suit the user's individual characteristics and needs. For example, when a user is learning a certain technique, the system can recommend other related techniques to help the user master vocal skills more comprehensively. By constructing a knowledge graph of vocal techniques and concrete metaphors, this invention not only provides users with richer and more intuitive guidance on vocal techniques but also achieves intelligent recommendation and correlation analysis of vocal techniques. This will help users better understand and master vocal techniques, improving vocal effects and expressiveness.

[0063] S403. The knowledge graph is input into a pre-trained diffusion model, and the diffusion model generates personalized training guidance text corresponding to vocal techniques based on the input knowledge graph. In step S403, it is important to explain in detail that the diffusion model is a probabilistic generative model that generates new samples by learning the latent distribution of data. In this embodiment, the diffusion model is used to generate personalized training guidance text corresponding to vocal techniques based on the input knowledge graph. To achieve this goal, a pre-trained diffusion model is first required. This model has already learned a large amount of text data related to vocal techniques during the training phase, thus possessing the ability to generate high-quality personalized training guidance text.

[0064] Before inputting the knowledge graph into the diffusion model, the system performs appropriate preprocessing. This includes extracting key information, constructing feature vectors, and performing necessary format conversions. The purpose of preprocessing is to transform the knowledge graph into an input format that the diffusion model can understand and process.

[0065] After preprocessing, the knowledge graph is input into a pre-trained diffusion model. Based on the input knowledge graph, the diffusion model progressively generates personalized training guidance text corresponding to vocal techniques. During generation, the model considers entity nodes (vocal techniques) and relational edges (figurative metaphors) in the knowledge graph, as well as the connections and hierarchical structure between them, to ensure that the generated text is both accurate and meets the user's needs.

[0066] The generated personalized training guidance texts can include detailed step-by-step descriptions, technique analyses, and practice suggestions, all designed to help users better understand and master vocal techniques. These texts not only contain scientific vocal principles and methods but also incorporate vivid metaphors and accompanying visual materials, making the learning process more intuitive and engaging.

[0067] S404. Perform semantic analysis and sentiment recognition on the personalized training guidance text, extract key information and sentiment tendencies, and select visual materials that match the vocal techniques and figurative metaphors from the preset visual material library based on the extracted key information and sentiment tendencies. In step S404, it is necessary to explain in detail that, in order to further improve the effectiveness of personalized training guidance, the system will perform in-depth semantic analysis and sentiment recognition on the generated personalized training guidance text. Semantic analysis aims to understand the key information and core points in the text, ensuring that users can accurately obtain the core guidance on vocal techniques. Sentiment recognition is used to capture the emotional tendency in the text, such as positive, negative, or neutral, which helps the system provide guidance that is more in line with the user's emotional needs.

[0068] During the semantic analysis phase, the system utilizes natural language processing (NLP) technology to parse the personalized training guidance text. Through steps such as word segmentation, part-of-speech tagging, and syntactic analysis, the system can identify key terms, phrases, and sentence structures in the text. This information is used to extract the core content of the text, ensuring that users can clearly understand the essentials and operational steps of each vocal technique.

[0069] The sentiment recognition stage relies on sentiment analysis algorithms (such as existing LSTM networks or BERT models) to determine the sentiment orientation of the text. This algorithm can identify sentimental words, expressions, and the overall sentiment tone within the text. Through sentiment analysis, the system can understand the emotional experiences users may have while learning vocal techniques. For example, when the text expresses a positive sentiment, the system can select inspiring visual materials to enhance the user's learning motivation; while when the text expresses a negative or confused sentiment, the system can provide additional explanations and demonstrations to help the user overcome difficulties.

[0070] After completing semantic analysis and sentiment recognition, the system selects visual materials from a pre-set visual material library that match vocal techniques and concrete metaphors based on the extracted key information and sentiment trends. These visual materials can be images, animations, videos, etc., and are designed to intuitively demonstrate the essentials and effects of vocal techniques. By combining these visual materials with personalized training guidance text, the system can provide users with richer, more vivid, and easier-to-understand learning resources.

[0071] S405. The selected visual materials are integrated with the personalized training guidance text to generate a concrete metaphorical description and accompanying visual materials.

[0072] In step S405, it is necessary to explain in detail that the system will deeply integrate the selected visual materials with the personalized training guidance text. This integration process aims to ensure a close connection and consistency between the visual materials and the text content. Specifically, the system will carefully select and edit the visual materials based on the vocal techniques and concrete metaphors described in the text. This ensures that each visual material accurately reflects the key points and essence of the vocal techniques while echoing the descriptions in the text. During the integration process, the system will also consider the user's cognitive habits and learning style. For example, for users who prefer intuitive learning, the system will use more visual materials such as images and animations to assist in the explanation; while for users who prefer in-depth thinking, the system may provide more detailed and in-depth textual explanations. Through this personalized integration method, the system can provide users with vocal technique guidance that is more tailored to their learning needs.

[0073] The final generated concrete metaphorical descriptions and accompanying visual materials will be presented to the user. These resources not only contain scientific principles and methods of vocalization, but also make the learning process more vivid and interesting through concrete metaphors and intuitive visual materials. Users can choose a learning method and pace that suits their own situation and needs, gradually mastering the essentials and application methods of vocalization techniques. This will help users improve their vocal effects and expressiveness, achieving more outstanding vocal performances.

[0074] As an optional embodiment of the present invention, optionally, the expression for generating personalized training guidance text corresponding to vocal techniques in step S403 is: in, Indicated based on knowledge graph Personalized training guidance text generation probability, The parameters represent the diffusion model. This refers to the final generated personalized training guidance text. A knowledge graph representing vocal techniques and metaphors. This represents the total time step of the diffusion process. Indicating vocal techniques, Indicates a relationship. It represents a concrete metaphor. Indicates the number of triples in the knowledge graph; The key information and sentiment expression extracted in step S04 are as follows: in, This indicates the source of personalized training guidance text. The set of keywords extracted from it. This represents the keyword extraction function. This refers to a terminology database specifically for the field of vocal music teaching. This represents the keyword confidence threshold. This represents a probability distribution vector representing sentiment tendencies. This indicates the activation of the normalization function. This represents the output layer weight matrix of the sentiment analysis model. Represents the Long Short-Term Memory network. This represents the bias term of the sentiment analysis model.

[0075] As an optional embodiment of the present invention, optionally, the expression for generating the recommended vocal techniques that adapt to the user's physiological conditions and artistic goals in step S5 is: in, This represents the final generated vector of vocal technique recommendations. Represents the normalized activation function. This represents the output layer weight matrix of the multimodal fusion model. This indicates a dimensional concatenation operation. Indicates the first The gating weights of each modality Indicates the first Feature vectors of each modality This represents the bias term of the output layer. The loss function represents multimodal fusion. Indicates the number of training samples. Indicates the first The true label of each sample The model represents the first The probability of a recommendation result for a sample. Represents the L2 regularization coefficient. This represents the squared L2 norm of the output layer weight matrix.

[0076] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for recommending vocal techniques based on deep learning, characterized in that, The method includes: S1. The original sound waves emitted by the user are hierarchically encoded through the brain-like acoustic perception layer to obtain acoustic feature vectors. S2. Construct a digital twin model of the user's vocal organs based on the user's physiological data, and generate personalized 3D biomechanical models of the vocal cords, larynx, and oral cavity through finite element analysis and fluid dynamics simulation. Simulate the biomechanical changes and acoustic effects during vocalization through a real-time physics engine. S3. Establish an adversarial aesthetic evaluation system, which includes a gold standard dataset and a public preference comparison learning framework. Through the multi-task public preference comparison learning framework, the acoustic rules of professional judges and the public's listening preferences are mapped to the same semantic space to generate a combination of vocal techniques that are both scientific and artistic. S4. Construct a cross-modal metaphor generation system based on vocal technique tags, establish a knowledge graph of vocal techniques and figurative metaphors based on natural language processing methods, generate personalized training guidance based on the knowledge graph through a diffusion model, and transform the abstract personalized training guidance into perceptible figurative metaphor descriptions and supporting visual materials. S5. The acoustic feature vectors, biomechanical changes, acoustic effects, vocal technique combinations, metaphorical descriptions, and accompanying visual materials are fused using a multimodal fusion algorithm to generate vocal technique recommendations that are adapted to the user's physiological conditions and artistic goals.

2. The deep learning-based vocal technique recommendation method as described in claim 1, characterized in that, The brain-like acoustic perception layer in step S1 includes a spiking neural network, which extracts the spatiotemporal pulse pattern, time-related features, and abstract acoustic features of sound waves through primary, intermediate, and advanced layers, respectively.

3. The deep learning-based vocal technique recommendation method as described in claim 2, characterized in that, The expression for the primary layer is: in, Indicates the first A neuron in time membrane potential, This represents the time step in a discrete-time simulation. This represents the membrane potential decay time constant. This represents the input current after sound wave conversion. Indicates the first Synaptic weights of individual neurons Represents neurons Pulse firing status, This represents the neuron firing threshold, where 1 indicates firing and 0 indicates resting. The expression for the intermediate layer is: in, Represents presynaptic neurons to postsynaptic neurons The change in connection weights, This represents the learning rate within the positive time window. Indicates the pulse timing difference. This represents the decay constant of the positive time window. This represents the decay constant for the negative time window. Represents time-related feature vectors. This indicates a pooling operation. Represents postsynaptic neurons The pulse state; The expression for the higher-level layer is: in, This represents the abstract acoustic feature vector output by the higher-level layer. Represents a non-linear activation function. Indicates the first Synaptic weights of individual neurons Indicates the first A neuron in the time window Weighted pulse count within, Indicates the first Bias terms for each neuron, Indicates the first A neuron in time The pulse state at that time, This represents the pulse count decay time constant.

4. The deep learning-based vocal technique recommendation method as described in claim 1, characterized in that, In step S2, the biomechanical changes and acoustic effects during vocalization are simulated using a real-time physics engine, including: S201. Preprocess the user's physiological data; S202. Based on the preprocessed physiological data of the user, establish a finite element model of vocal cord biomechanics, a finite element model of laryngeal fluid, and a finite element model of oral cavity tuning; S203. The vocal cord biomechanical finite element model, the laryngeal cavity fluid finite element model and the oral cavity tuning finite element model are bidirectionally coupled to obtain a digital twin model of the user's vocal organs. S204. Import the digital twin model into the real-time physics engine and optimize the model parameters of the digital twin model using the Bayesian optimization algorithm. S205. Adjust the vocal cord biomechanical finite element model in the digital twin model using a non-rigid registration algorithm; optimize the laryngeal cavity cross-sectional area change curve using airflow pressure data from the user's physiological data, and define a personalized relationship between subglottic pressure and glottal impedance; optimize the tongue trajectory and lip shape parameters in the oral cavity tuning finite element model based on vocalization task data using a genetic algorithm. S206. Based on the optimized model parameters of the digital twin model, the biomechanical changes and acoustic effects during sound production are simulated using finite element analysis and fluid dynamics simulation methods.

5. The deep learning-based vocal technique recommendation method as described in claim 4, characterized in that, The expression for the non-rigid registration algorithm in step S205 is: in, This represents the optimal displacement field. Indicates the number of corresponding point pairs. Indicates the first The spatial location of each coordinate point Indicates the first The original positions of the coordinate points Represents the displacement field function. This represents the regularization weight coefficient. Indicates the number of nodes in the target model. This represents the second-order gradient norm of the displacement field. Indicates the first The original positions of the coordinate points This represents the weight coefficient of the constraint term. The first part of the vocal cord after deformation One morphological feature, This represents the first [number]th ... Reference values ​​for morphological features.

6. The deep learning-based vocal technique recommendation method as described in claim 1, characterized in that, The combination of vocal techniques that combines scientific rigor and artistic expression generated in step S3 includes: S301. Based on professional acoustic principles and public listening preference data, construct the gold standard dataset and the public preference dataset; S302. Utilize a pre-trained acoustic model as a shared encoder, and extract professional acoustic features and public perception features based on the gold standard dataset and the public preference dataset, and output scientific and artistic scores respectively through a dual-branch network. S303. Based on the scientific and artistic scores, an adversarial training strategy is used to distinguish between the professional feature distribution and the general feature distribution using a discriminator, thereby prompting the shared encoder to generate unified professional and general features that are independent of modality. S304. By comparing and learning, the semantic distance between the professional features and the general features is narrowed, and by combining dynamic weighting with gating units, a fusion feature that combines both evaluation criteria is obtained. S305, Based on the fusion features, the decoder is used to generate a combination of vocal techniques.

7. The deep learning-based vocal technique recommendation method as described in claim 6, characterized in that, The expression for the adversarial training strategy in step S303 is: in, Represents the adversarial training loss function. This represents a sample from the gold standard dataset. Indicates the distribution of professional data. Represents the mathematical expectation. Indicates the discriminator, Indicates a shared encoder. This represents a sample of a mass preference dataset. Indicates the distribution of mass data; The expression for obtaining the fusion feature that combines both evaluation criteria in step S304 is: in, This represents the fused feature vector. This represents the Sigmoid activation function. This represents the weight matrix of the gated unit. Represents professional acoustic feature vectors. Represents the feature vector of public perception. Indicates the gate unit bias term. This represents element-wise multiplication. This represents the contrastive learning semantic alignment loss. Represents the cosine similarity function. Indicates the temperature coefficient. Indicates the number of negative samples. Indicates the first The public's perceived feature vector of a negative sample This represents the weighting coefficient for the comparative loss.

8. The deep learning-based vocal technique recommendation method as described in claim 1, characterized in that, In step S4, the abstract personalized training guidance is transformed into a perceptible, concrete metaphorical description and accompanying visual materials, including: S401. Using natural language processing methods, extract keywords of vocal techniques and their corresponding concrete metaphorical descriptions from vocal music teaching literature, and construct a database of the correspondence between vocal techniques and concrete metaphors. S402. Based on the aforementioned correspondence database, a knowledge graph of vocal techniques and figurative metaphors is constructed using a knowledge graph construction method, wherein vocal techniques are entity nodes in the knowledge graph, and figurative metaphors are relation edges associated with the entity nodes of vocal techniques. S403. The knowledge graph is input into a pre-trained diffusion model, and the diffusion model generates personalized training guidance text corresponding to vocal techniques based on the input knowledge graph. S404. Perform semantic analysis and sentiment recognition on the personalized training guidance text, extract key information and sentiment tendencies, and select visual materials that match the vocal techniques and figurative metaphors from the preset visual material library based on the extracted key information and sentiment tendencies. S405. The selected visual materials are integrated with the personalized training guidance text to generate a concrete metaphorical description and accompanying visual materials.

9. The deep learning-based vocal technique recommendation method as described in claim 8, characterized in that, The expression for generating personalized training guidance text corresponding to vocal techniques in step S403 is as follows: in, Indicated based on knowledge graph Personalized training guidance text generation probability, The parameters represent the diffusion model. This refers to the final generated personalized training guidance text. A knowledge graph representing vocal techniques and metaphors. This represents the total time step of the diffusion process. Indicating vocal techniques, Indicates a relationship. It represents a concrete metaphor. Indicates the number of triples in the knowledge graph; The key information and sentiment expression extracted in step S04 are as follows: in, This indicates the source of personalized training guidance text. The set of keywords extracted from it. This represents the keyword extraction function. This represents a specialized terminology database for vocal music teaching. This represents the keyword confidence threshold. This represents a probability distribution vector representing sentiment tendencies. This indicates the activation of the normalization function. This represents the output layer weight matrix of the sentiment analysis model. Represents the Long Short-Term Memory network. This represents the bias term of the sentiment analysis model.

10. The deep learning-based vocal technique recommendation method as described in claim 1, characterized in that, The expression for generating the vocal technique recommendation result that adapts to the user's physiological conditions and artistic goals in step S5 is as follows: in, This represents the final generated vector of vocal technique recommendations. Represents the normalized activation function. This represents the output layer weight matrix of the multimodal fusion model. This indicates a dimensional concatenation operation. Indicates the first The gating weights of each modality Indicates the first Feature vectors of each modality This represents the bias term of the output layer. The loss function represents multimodal fusion. Indicates the number of training samples. Indicates the first The true label of each sample The model represents the first The probability of a recommendation result for a sample. Represents the L2 regularization coefficient. This represents the squared L2 norm of the output layer weight matrix.