Human voice customization simulation and song generation method and system
By extracting multidimensional features from singers' vocal data to construct a four-dimensional mapping model, the problems of subjective and single-dimensional evaluation in traditional song synthesis technology are solved, generating songs that conform to personalized timbre and emotion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-03-31
AI Technical Summary
Existing song synthesis technologies lack objective, multi-dimensional quantitative indicators, have a single evaluation dimension, cannot fully reflect timbre quality, and are difficult to support AI timbre simulation and personalized song generation.
By acquiring vocal data and evaluation indicators from different singers, features from the acoustic, physiological, aesthetic, and recognition layers are extracted to construct a four-dimensional mapping model. This model is then trained using a positive correlation module and a negative feedback module to output the target human voice.
It achieves a complete depiction from physical foundations to artistic expression, and the generated music not only conforms to physical laws but also meets aesthetic needs. The personalized timbre is highly synchronized with the lyrics' emotions and the melody's structure.
Smart Images

Figure CN121768359A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of audio processing technology, and in particular to a method and system for customized human voice simulation and song generation. Background Technology
[0002] With the continuous advancement of technology, audio processing technology is also constantly developing. However, as user needs continue to change, some consumer electronics and entertainment applications not only want to synthesize speech but also songs, which necessitates training a model with song synthesis capabilities.
[0003] However, existing song synthesis technologies have the following problems: (1) Traditional timbre evaluation relies on subjective feelings and lacks objective, multi-dimensional quantitative indicators, resulting in poor consistency of evaluation results; (2) The evaluation dimensions are too singular to fully reflect the quality of tone; (3) It is difficult to support the needs of AI timbre simulation, personalized song generation and other scenarios, and there is a lack of structured scoring data that can be used for model training or optimization.
[0004] Therefore, there is an urgent need to develop a method and system for customized human voice simulation and song generation to solve the above problems. Summary of the Invention
[0005] This disclosure provides a method and system for customized vocal simulation and song generation, which solves the problem that existing song synthesis systems lack objective and multi-dimensional quantitative indicators, have a single evaluation dimension, and cannot fully reflect the timbre quality.
[0006] According to a first aspect of this disclosure, a method and system for customized vocal simulation and song generation are provided. The method includes: acquiring vocal data, vocal text, and evaluation metrics from different singers as a sample set, and preprocessing the sample set; The preprocessed sample set is subjected to feature extraction to obtain four-dimensional features and contextual features. The four-dimensional features include acoustic layer features, physiological layer features, recognition layer features and aesthetic layer features. The four-element mapping model is trained using the four-element features and contextual features. The four-element mapping model is used to output acoustic signals based on the correlation between the four-element features. Input the target singer's vocal data to be processed into the trained four-image mapping model, and output the target vocal data.
[0007] In addition to the aspects described above and any possible implementations, a further implementation is provided in which the four-element mapping model is constructed based on the following formula: , Where V represents the output sound. This represents the acoustic layer characteristics at time t. (·) indicates acoustic layer characteristics Physiological layer characteristics at time t Aesthetic layer features A, and recognition layer features at time t. And a function of contextual features C.
[0008] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the four-element mapping model includes an input layer, an association layer, and an output layer, wherein the association layer includes a positive association module and a negative feedback module; The positive correlation module is constructed based on the positive correlation between the four features, and the negative feedback module is constructed based on the negative correlation between the four features.
[0009] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the positive correlation includes: obtaining acoustic layer features by driving the output acoustic signal based on the physiological layer features; obtaining acoustic layer features based on the constraints of the recognition layer features; and using the aesthetic layer features as an optimization function for the quality of the acoustic layer features. The reverse correlation includes: using aesthetic layer features as an optimization function to correct physiological layer features, and using the output acoustic signal to correct physiological layer features.
[0010] In addition to the aspects described above and any possible implementations, a further implementation is provided in which the positive correlation module can be represented by the following formula: Where B represents physiological features, S represents acoustic features, A represents aesthetic features, and I represents recognition features. This represents the physiological layer characteristics after they are interconnected. This represents the acoustic layer features after they are correlated with each other. This represents the aesthetic features that are interconnected. R represents the recognition layer features after mutual correlation, C represents the four-dimensional influence matrix, and C represents the contextual features. In the four-element influence matrix, the elements at different positions represent different positive correlation values between the four elements.
[0011] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the reverse feedback module includes the following steps: Calculate the acoustic layer error and aesthetic layer error based on the currently output acoustic signal; Determine whether the acoustic layer error does not exceed the recognition layer tolerance and the aesthetic layer error does not exceed a preset difference; if not, calculate the correction intensity. The current physiological layer features are corrected according to the correction intensity to obtain the updated physiological layer features; Based on the updated physiological layer characteristics, the acoustic signal is re-output, and the aforementioned steps are repeated until the acoustic layer error does not exceed the recognition layer tolerance and the aesthetic layer error does not exceed the preset difference.
[0012] In addition to the aspects and any possible implementations described above, a further implementation is provided, which further includes: The current emotional impact can be calculated using the following formula: in, Indicates the current physiological depth. Indicates aesthetic matching, Indicates current recognition memory. Indicates current emotional vocal energy. Indicates the degree of contextual matching. This indicates the updated acoustic layer error. This represents the overall strength or impact force weighting coefficient; Determine if the current emotional impact is less than the target value. If so, continue adjusting the physiological characteristics until the updated emotional impact is less than the target value.
[0013] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the vocal data of the target singer to be processed is input into the trained four-dimensional mapping model, and the target vocal data is output, including: Extract the four-dimensional features and contextual features of the vocal data to be processed; The four-dimensional features to be processed are adjusted using the context features to be processed, resulting in fused four-dimensional features. The fused four-image features and the context features to be processed are input into the trained four-image mapping model, and the initial human voice is output. The target voice is obtained by performing reverse optimization based on the initial human voice through a reverse feedback module.
[0014] In addition to the aspects and any possible implementations described above, a further implementation is provided, which adjusts the four-dimensional features to be processed using the contextual features to be processed to obtain fused four-dimensional features, including: Each sub-feature in the context features to be processed is respectively constrained to the four-image features to obtain the fused four-image features.
[0015] According to a second aspect of this disclosure, a customized vocal simulation and song generation system is provided, comprising: a data acquisition module: acquiring vocal data, vocal text, and evaluation indicators of different singers as a sample set, and preprocessing the sample set; Feature extraction module: connected to the data acquisition module, used to extract features from the preprocessed sample set to obtain four-dimensional features and contextual features. The four-dimensional features include acoustic layer features, physiological layer features, recognition layer features and aesthetic layer features. Model training module: connected to the feature extraction module, used to train the four-element mapping model using the four-element features and contextual features, the four-element mapping model is used to output acoustic signals based on the correlation between the four-element features; Music output module: connected to the model training module, used to input the vocal data of the target singer to be processed into the trained four-dimensional mapping model and output the target vocal data.
[0016] The beneficial effects of this disclosure are: This disclosure extracts features from four dimensions: acoustic, physiological, aesthetic, and recognition, which are used to capture the objective physical properties of sound, reveal the physiological control mechanism of sound generation, reflect aesthetic value, and identify individual style. It breaks through the limitations of traditional single-dimensional focus on pitch, volume, frequency, etc., and realizes the full-process depiction from physical basis to artistic expression, fully reflecting the essence of timbre. It also solves the problems of subjectivity and ambiguity in traditional evaluation. A positive correlation module was set up based on the interrelationships and mutual influences between the acoustic layer, physiological layer, aesthetic layer, and recognition layer, which avoids the problem of timbre being disconnected from physiological mechanisms in traditional models; This invention incorporates a reverse feedback module to correct deviations in real time, ensuring that the output conforms to both physical laws and aesthetic requirements; the final generated music not only matches the personality of the target timbre but also is highly synchronized with the lyrics' emotions and the melody's structure.
[0017] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0018] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A flowchart of a method for customized human voice simulation and song generation provided in an embodiment of this disclosure is shown; Figure 2 An association diagram provided in an embodiment of this disclosure is shown; Figure 3A framework diagram of the voice customization simulation and song generation system provided in this disclosure embodiment is shown; Figure 4 A block diagram of an electronic device provided according to an embodiment of the present disclosure is shown. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0020] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0021] This disclosure provides a method for customized vocal simulation and song generation. See [link to relevant documentation]. Figure 1 This includes the following steps: S1. Obtain vocal data, vocal texts, and evaluation indicators from different singers as a sample set, and preprocess the sample set.
[0022] The vocal data should include singing excerpts of different styles and emotions performed by the singers, with each excerpt ideally longer than 30 seconds to improve the accuracy of the subsequent model training. The vocal text refers to the corresponding lyrics, and the evaluation metrics are expert ratings or audience perception questionnaire results for different singers' performance levels, primarily focusing on subjective indicators such as "pleasure" and "emotional delivery." Preprocessing includes data cleaning, outlier detection, and duplicate value removal.
[0023] S2. Extract features from the preprocessed sample set to obtain four-dimensional features and contextual features. The four-dimensional features include acoustic layer features, physiological layer features, recognition layer features and aesthetic layer features. This publication constructs a comprehensive vocal analysis and generation system based on four core dimensions. These four dimensions encompass the four essential dimensions of the human voice, which are interconnected and mutually influential, forming a multidimensional vocal thinking system that can both describe and operate: (1) The acoustic layer is the physical manifestation of sound energy. It is a measurable external sound phenomenon, including frequency, overtones, envelope, HNR, density, loudness, etc.
[0024] (2) The physiological layer is the sound-producing mechanism of human organs and muscles, which is the root of sound source control, including vocal cords, muscles, resonance cavities, airflow, nerves, etc.
[0025] (3) The aesthetic layer represents human aesthetic perception and evaluation of sound. It is the aesthetic lens of the audience and the producer. It is the standard of "good", "bad" and "moving", including such as range, timbre, expressiveness, emotional depth and control.
[0026] (4) The recognition layer represents the style and personality of the voice, records all the patterns of the voice, and is the answer to "who is singing" and "why it is different from others", such as pronunciation, melody habits, intonation, narrative tendency, rhythm preference, etc.
[0027] The four elements influence and are interconnected, exhibiting the following relationships: (1) For example, the physiological layer muscle regulation can change the acoustic layer spectral behavior.
[0028] (2) For example, the glottal opening and closing pattern and habitual muscle pathways in the physiological layer can affect articulation and intonation style.
[0029] (3) For example, the frequency distribution of the acoustic layer will bring a “pleasant” and “warm” listening experience to the aesthetic layer.
[0030] (4) For example, the style preferences of singers in the identification layer will also affect the aesthetic evaluation of the audience in the aesthetic layer.
[0031] (5) For example, deliberately training muscles / control to achieve a certain aesthetic, that is, reverse control from the aesthetic level to the physiological level.
[0032] (6) For example, sound features enhance personalized expression, which in turn are shaped by style, that is, the acoustic layer and the recognition layer influence each other.
[0033] Acoustic layer features can be extracted using audio analysis tools (such as MATLAB, Librosa, etc.). By calculating the FFT spectrum, frequency distribution, overtone structure, formant positions, envelope ADSR parameters, etc., in vocal data can be obtained. Energy density can also be calculated using RMS. These parameters are all acoustic layer features, accurately capturing the physical properties of sound and used to evaluate sound thickness and dynamic changes, providing a foundation for subsequent analysis of physiological and aesthetic layers.
[0034] For physiological layer characteristics, acoustic layer characteristics can be used to infer their effects using physical laws. For example, vocal cord length can be inferred from the fundamental frequency of the acoustic layer (the longer the vocal cords, the lower the frequency); vocal cord tension can be inferred from high-frequency brightness (the higher the vocal cord tension, the more pronounced the high frequencies); respiratory efficiency can be analyzed through airflow stability; and vital capacity can be inferred from the ADSR (Advanced Surface Range) parameters of the acoustic layer (the longer the release time in the ADSR parameters, the greater the vital capacity). By analyzing physiological layer characteristics, the root causes of sound production can be revealed, establishing a correlation between physiological actions and acoustic layer performance, providing a foundation for subsequent reverse training.
[0035] Regarding aesthetic features, this disclosure requires the extraction of evaluation indicators for each singer. Emotional transmission power can be evaluated by constructing an emotion recognition model, which includes spectrograms and semantic analysis. Pitch control power can be calculated through pitch offset values, which can quantify various evaluations and provide data basis for subsequent style optimization.
[0036] Regarding the identification layer features, this disclosure can extract pronunciation habits, such as the trailing tone pattern of specific syllables, through voiceprint fingerprints; match decorative tone deviations based on melody variation templates, such as the frequency of use of glissando and vibrato; and analyze personal rhythm processing style through rhythm offset rate, etc.
[0037] Physiological parameters include effective data on the singer's physiological structure, such as vocal cord length, collected through tension sensors; aesthetic annotations include expert scores on the singer's vocal range and emotional delivery, serving as aesthetic evaluation indicators; style tags are annotations on the singer's enunciation habits, rhythm preferences, etc.
[0038] Contextual features include song style, emotional tone, tempo, and lyrics, which can be inferred from the four key characteristics. For example, a deep, low-frequency sound can indicate classical music, a bright, high-frequency sound can indicate pop, and clear, concise enunciation can indicate a fast-paced song.
[0039] S3. The four-symbol mapping model is trained using the four-symbol features and contextual features. The four-symbol mapping model is used to output acoustic signals based on the correlation between the four-symbol features.
[0040] The four-element mapping model includes an input layer, an association layer, and an output layer. The association layer includes a positive association module and a negative feedback module. The positive association module is constructed based on the positive association relationship between the four-element features, and the negative feedback module is constructed based on the negative association relationship between the four-element features.
[0041] The four-element mapping model is constructed based on the following formula: , Where V represents the output sound. This represents the acoustic layer characteristics at time t, derived from the physiological layer characteristics at time t. Aesthetic layer features A, and recognition layer features at time t. And contextual feature C together determine, (·) indicates acoustic layer characteristics Regarding physiological layer characteristics Aesthetic layer features A, Recognition layer features And a function of contextual features C.
[0042] Among them, the recognition layer features need to be updated in real time according to the singing status: , express The features of the recognition layer at each moment. Represents style increment function, physiological layer action The aesthetic layer A and the context C will evolve during the performance, forming a style increment function. The style increment of the previous moment will affect the next moment. Therefore, it is necessary to embed the style increment into the style memory to reflect the latest pronunciation habits, rhythm preferences, etc.
[0043] In practice, the inverse short-time Fourier transform can be used to... It is converted into an acoustic signal for output.
[0044] In one specific embodiment, the positive correlation module can be represented by the following formula: Where B represents physiological features, S represents acoustic features, A represents aesthetic features, and I represents recognition features. This represents the physiological layer characteristics after they are interconnected. This represents the acoustic layer features after they are correlated with each other. This represents the aesthetic features that are interconnected. Let R represent the features of the recognition layer after they are correlated with each other, R represent the four-dimensional influence matrix, and C represent the contextual features. This formula shows the dynamic dependency and feedback relationship between the correlation layers over time.
[0045] See Figure 2The positive correlation among the four elements includes: obtaining acoustic layer features by driving the output acoustic signal based on the physiological layer features; obtaining acoustic layer features based on the constraints of the recognition layer features; and using the aesthetic layer features as an optimization function for the quality of the acoustic layer features. Specifically, the physiological layer is first used as the driving source, based on analytical parameters such as vocal cord tension and breathing efficiency, to drive the generation of the acoustic layer spectrum. Then, the style features of the recognition layer (such as intonation habits) constrain the output of the acoustic layer to achieve dynamic changes. Finally, the aesthetic layer objective (such as whether the emotional depth meets the standard) serves as the optimization function for the output quality of the acoustic layer.
[0046] The elements at different positions in the four-element influence matrix represent different positive correlation values between the four elements, which can be preset according to actual conditions. For example, physiological layer features drive the acquisition of acoustic layer features, which serve as the main path; the parameters at the corresponding positions in the four-element influence matrix can be set as follows. Physiological layer features also determine recognition layer features. As a secondary path, the parameters at the corresponding positions in the four-dimensional influence matrix are set as follows: Acoustic layer features determine aesthetic layer features; the two are related, with the former being the foundation of the latter. The parameters at the corresponding positions in the four-dimensional influence matrix are set as follows: The features of the recognition layer influence the features of the aesthetic layer, which belongs to the style and aesthetic judgment path. The parameters at the corresponding positions in the four-element influence matrix can be set as follows: .
[0047] In the statistical model, the values of each element in the four-dimensional influence matrix are in the range of [-1, 1], in order to analyze the correlation between various variables, such as the influence of physiological characteristics on acoustic characteristics. It is necessary to ensure that the changes in vocal cord tension in physiological characteristics are linearly correlated with the changes in high-frequency overtones in acoustic characteristics.
[0048] In addition, the following reverse correlations also exist: The aesthetic layer features serve as the final target, and the optimization function can be used to train the physiological layer features in reverse. Based on the actual output acoustic layer physical signal, the physiological layer features are corrected. The acoustic layer features and the recognition layer features influence each other.
[0049] Based on the reverse correlation, this disclosure sets up a reverse feedback module to update the four-element influence matrix in real time according to the output results, and to update the correlation between the four elements by updating the values of each element.
[0050] Specifically, the reverse feedback module includes: S31. Calculate the acoustic layer error and aesthetic layer error based on the current output acoustic signal, quantify the error, locate the gap between the current output and the ideal situation, provide accurate basis for subsequent correction feedback, and avoid blind adjustment.
[0051] By calculating acoustic layer error : in, Indicates the expected sound to be emitted. This represents the currently output acoustic signal. This indicates the current moment. The calculation principles for aesthetic layer error, recognition layer error, and acoustic layer error are similar, and will not be elaborated here. Among them, acoustic layer error directly reflects the deviation of sound physical characteristics, while aesthetic layer error and recognition layer error reflect perceptual and stylistic deviations.
[0052] S32. Determine whether the acoustic layer error does not exceed the recognition layer tolerance and the aesthetic layer error does not exceed the preset difference; if not, calculate the correction intensity; if yes, no correction is needed.
[0053] Specifically, if and , This indicates the preset recognition layer tolerance. This indicates the aesthetic layer error at the current moment. The identification layer tolerance and preset difference can be set according to the actual situation. If the aforementioned conditions are met, it means that the acoustic layer error is within the range allowed by the style and does not need to be corrected.
[0054] like and This indicates that an offset has occurred, so the correction strength is calculated according to the following formula. : in, This indicates the self-control ability at the current moment. The larger the value, the stronger the self-control ability. It can be inferred from physiological characteristics, such as the range of muscle control fluctuations. The smaller the fluctuation range, the greater the self-control ability. This represents the fatigue coefficient at the current moment. The larger the value, the more severe the fatigue and the weaker the correction ability. It can also be obtained through physiological layer characteristic analysis, such as the degree of decline in respiratory efficiency. The greater the decline, the larger the value. This indicates that a smoothing operation is performed to avoid excessive correction strength due to sudden error changes, which would result in an invalid correction strength.
[0055] In summary, by combining stylistic characteristics and one's own control, we can avoid over-correction or ineffective correction, and make the correction more in line with human singing logic.
[0056] S33. Correct the current physiological layer features according to the correction intensity to obtain the updated physiological layer features.
[0057] Let the physiological characteristics at the current moment be... The updated physiological layer features Calculated using the following formula: in, It can represent the moment after the current moment, or the moment after the current moment's data has been updated. The increment of physiological layer features is represented by the following formula: in, This indicates the preset action plan at the current moment. For example, raising the pitch requires shortening the vocal cords and increasing muscle tension. This value is determined by the mapping rules between acoustics and aesthetics. The mapping relationship can be organized into a table as the basis for the preset action plan. (·)express The mapping relationship between them.
[0058] In summary, by adjusting the matching error type through physiological parameters and being constrained by physiological limits, damage can be avoided while ensuring the reduction of errors.
[0059] S34. Based on the updated physiological layer characteristics, output the acoustic signal again, repeat S31 to S33 until the acoustic layer error does not exceed the recognition layer tolerance and the aesthetic layer error does not exceed the preset difference, and then output the corresponding acoustic signal at this time.
[0060] The re-output acoustic signal is represented by the following formula: , in, For the updated recognition layer features, specifically through , in, This represents the recognition layer features at the current moment.
[0061] In summary, by judging the error of the output acoustic signal, correcting the physiological layer characteristics, and optimizing the acoustic layer, aesthetic layer, and recognition layer in a coordinated manner, the most suitable acoustic signal is output to meet the aesthetic requirements of being pleasant to listen to and moving people.
[0062] Furthermore, this disclosure also includes content regarding the core stimulation of emotions. Emotional impact is a core objective of the aesthetic layer, achieved by strengthening physiological control and reducing errors to meet the requirements of the aesthetic layer. The current emotional impact is first calculated using the following formula: , in, It indicates the current physiological depth to quantify the degree to which the characteristics of the current physiological layer support emotional expression. For example, in the case of a calm emotion, the physiological movements are relatively relaxed. This can be obtained by multi-dimensional reasoning and calculation of the physiological layer corresponding to the acoustic signal. Aesthetic matching refers to the degree of conformity between the output acoustic signal and the preset aesthetic goal. Different aesthetic index scoring standards can be set, and the aesthetic matching value corresponding to the acoustic signal can be calculated according to the standard. This indicates the current recognition memory, which is the degree to which the style of the currently output acoustic signal matches the characteristics of the corresponding singer's style. It can be obtained through feature analysis and comparison of the recognition layer. Representing current emotional acoustic energy involves quantifying the physical properties of acoustic signals into emotion-related quantities. The context matching degree represents the degree to which the emotional expression of the current output acoustic signal matches the context of the song. It can be calculated through multi-dimensional adaptation analysis. This indicates the updated acoustic layer error. This represents the overall intensity or impact weighting coefficient, and different default values can be set for different music scenarios.
[0063] The above formula includes an error suppression term. : The larger the value, the smaller the time error suppression term, meaning the closer the actual vocalization is to the target, and the stronger the emotional impact; a confidence adjustment factor has also been added. The more experience one has, the faster and stronger the feedback and adjustment will be.
[0064] Determine if the current emotional impact is less than the target value. If so, continue to adjust the physiological characteristics until the updated emotional impact is less than the target value.
[0065] S4. Input the vocal data of the target singer to be processed into the trained four-image mapping model, and output the target vocal data.
[0066] S41. Extract the four-dimensional features and contextual features of the vocal data to be processed, so as to provide a basis for generating suitable songs in the future.
[0067] The vocal data to be processed should include different singing segments at different pitches and intensities, with each segment preferably longer than 30 seconds. Specific methods for feature extraction can be found in section S2 above, and will not be elaborated upon here.
[0068] The contextual features to be processed include not only the song style and emotional tone, but also the lyrics and melody. Clarifying the contextual features to be processed is to clarify the external constraints of the song and ensure that the final output song matches the lyrics.
[0069] S42. Adjust the four-image features to be processed using the context features to be processed to obtain the fused four-image features; Each sub-feature in the context features to be processed is respectively constrained to the four-image features to obtain the fused four-image features.
[0070] Sub-features refer to the various features included in the context features to be processed. Each feature corresponds to one of the four-element features. The context features to be processed restrict the range of the four-element features to avoid the subsequent output songs not meeting the requirements.
[0071] For songs like rap, a fast tempo requires a spectral flux of 10Hz or higher per frame, while an intense emotional expression requires a high level of emotional delivery, thus meeting the basic requirements.
[0072] S43. Input the fused four-image features and the context features to be processed into the trained four-image mapping model, and output the initial human voice; By fusing the four-dimensional features and the context features to be processed, the positive correlation module is activated. Finally, the acoustic layer feature sequence can be converted into the target human voice output through the InverseSTFT algorithm.
[0073] S44. Based on the initial human voice, reverse optimization is performed through the reverse feedback module to obtain the target human voice.
[0074] By setting a reverse feedback module to modify the parameters of the forward correlation module, the output can be adjusted, ultimately outputting the target vocals corresponding to the target singer under the set lyrics and melody.
[0075] This disclosure can also achieve the purpose of directly modifying the timbre state by directly modifying the four-element influence matrix or the four-element features to be processed in the positive correlation module.
[0076] Furthermore, this disclosure can score the timbre of a target singer, specifically including the following steps: quantifying the four characteristics to be processed according to a preset scoring standard; for example, the maximum score for each type of characteristic is 10 points; calculating and outputting the total score of the four characteristics to be processed. The final score reflects both the physical characteristics of the timbre and its aesthetic value and individual characteristics, providing a clear direction for timbre optimization or customization.
[0077] Based on the above technical solution, this disclosure extracts features from four dimensions: acoustic layer, physiological layer, aesthetic layer, and recognition layer. These features are used to capture the objective physical properties of sound, reveal the physiological control mechanism of sound generation, reflect aesthetic value, and identify individual style. It breaks through the limitations of traditional single-dimensional focus on pitch, volume, frequency, etc., and realizes the full-process depiction from physical basis to artistic expression, fully reflecting the essence of timbre. It also solves the problems of subjectivity and ambiguity in traditional evaluation. A positive correlation module was set up based on the interrelationships and mutual influences between the acoustic layer, physiological layer, aesthetic layer, and recognition layer, which avoids the problem of timbre being disconnected from physiological mechanisms in traditional models; This invention incorporates a reverse feedback module and corrects deviations in real time to ensure that the output conforms to both physical laws (such as the matching of vocal cord tension and pitch) and aesthetic requirements (such as achieving the required emotional delivery). The final generated music not only matches the personality of the target timbre but also is highly synchronized with the lyrics' emotions and the melody's structure.
[0078] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this disclosure.
[0079] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0080] The above is a description of the method embodiments. This disclosure also provides a customized human voice simulation and song generation system 300. See also... Figure 3 It includes the following modules: Data acquisition module 301: Acquires vocal data, vocal texts, and evaluation indicators of different singers as a sample set, and preprocesses the sample set; Feature extraction module 302: connected to the data acquisition module 301, used to extract features from the preprocessed sample set to obtain four-dimensional features and contextual features. The four-dimensional features include acoustic layer features, physiological layer features, recognition layer features and aesthetic layer features. Model training module 303: connected to the feature extraction module 302, used to train the four-element mapping model using the four-element features and contextual features, the four-element mapping model is used to output acoustic signals based on the correlation between the four-element features; Music output module 304: connected to the model training module 303, used to input the vocal data of the target singer to be processed into the trained four-image mapping model and output the target vocal data.
[0081] Other details can be found in the previous methods section and will not be repeated here.
[0082] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0083] Figure 4A schematic block diagram of an electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0084] Electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in ROM 402 or a computer program loaded into RAM 403 from storage unit 408. RAM 403 can also store various programs and data required for the operation of electronic device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. I / O interface 405 is also connected to bus 404.
[0085] Multiple components in electronic device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of displays, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0086] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the customized voice simulation and song generation method. For example, in some embodiments, the customized voice simulation and song generation method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of the customized voice simulation and song generation method described above can be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform the vocal customization simulation and song generation method by any other suitable means (e.g., by means of firmware).
[0087] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0088] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0089] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0090] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).
[0091] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0092] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0093] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0094] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for human voice customization simulation and song generation, characterized by, The method comprises: acquiring vocal data, vocal text and evaluation indexes of different singers as a sample set, and preprocessing the sample set; extracting features from the preprocessed sample set to obtain four-character features and context features, wherein the four-character features include acoustic layer features, physiological layer features, recognition layer features and aesthetic layer features; training a four-character mapping model using the four-character features and the context features, wherein the four-character mapping model is used to output acoustic signals based on the correlation between the four-character features; inputting processed vocal data of a target singer into the trained four-character mapping model to output target vocals.
2. The method of claim 1, wherein the four-character mapping model is constructed based on the following formula: The four-character mapping model comprises an input layer, a correlation layer and an output layer, wherein the correlation layer comprises a forward correlation module and a backward feedback module; , where V is the output sound, represents the acoustic layer feature at time t, (·) represents the acoustic layer feature is a function of the physiological layer feature at time t , the aesthetic layer feature A, the recognition layer feature at time t and the context feature C.
3. The method of customizing vocal simulation and song generation of claim 2, wherein, The forward correlation module is constructed based on the forward correlation between the four-character features, and the backward feedback module is constructed based on the backward correlation between the four-character features. The forward correlation includes: driving the output acoustic signal to obtain the acoustic layer features based on the physiological layer features; constraining the acoustic layer features based on the recognition layer features; and using the aesthetic layer features as an optimization function of the quality of the acoustic layer features; 4. The method of customizing vocal simulation and song generation according to claim 3, wherein, The backward correlation includes: using the aesthetic layer features as an optimization function to feedback and correct the physiological layer features, and using the output acoustic signal to feedback and correct the physiological layer features. The forward correlation module can be represented by the following formula:
5. The method of customizing vocal simulation and song generation of claim 4, wherein, wherein the elements at different positions in the four-character influence matrix represent different forward correlation values between the four characters. wherein B represents a physiological layer feature, S represents an acoustic layer feature, A represents an aesthetic layer feature, I represents an identification layer feature, represents a physiological layer feature after being associated with each other, represents an acoustic layer feature after being associated with each other, represents an aesthetic layer feature after being associated with each other, represents an identification layer feature after being associated with each other, R represents a four-influence matrix, and C represents a context feature. The backward feedback module comprises the following steps:
6. The method of customizing vocal simulation and song generation of claim 4, wherein, calculating acoustic layer errors and aesthetic layer errors based on the current output acoustic signal; determining whether the acoustic layer errors are less than the recognition layer tolerance and the aesthetic layer errors are less than a preset difference value; if not, calculating a correction strength; correcting the current physiological layer features based on the correction strength to obtain updated physiological layer features; reoutputting the acoustic signal based on the updated physiological layer features, and repeating the above steps until the acoustic layer errors are less than the recognition layer tolerance and the aesthetic layer errors are less than the preset difference value. Further comprising:
7. The method of customizing vocal simulation and song generation of claim 6, wherein, calculating the current emotional impact by the following formula: determining whether the current emotional impact is less than a target value; if so, adjusting the physiological layer features until the updated emotional impact is less than the target value. wherein, represents the current physiological depth, represents the aesthetic match, represents the current recognition memory, represents the current emotional sound energy, represents the contextual match, represents the updated acoustic layer error, represents the overall intensity or impact force weight factor; Inputting the processed vocal data of the target singer into the trained four-character mapping model to output target vocals, comprising:
8. The method of customizing vocal simulation and song generation of claim 3, wherein, extracting processed four-character features and processed context features from the processed vocal data; adjusting the processed four-character features using the processed context features to obtain fused four-character features; inputting the fused four-character features and the processed context features into the trained four-character mapping model to output initial vocals; performing backward optimization on the initial vocals based on the initial vocals to obtain target vocals. Adjusting the processed four-character features using the processed context features to obtain fused four-character features, comprising:
9. The method of customizing vocal simulation and song generation of claim 1, wherein, Corresponding to the four-character features, each of the sub-features in the context features to be processed is limited to obtain fused four-character features.
10. A human voice customisation simulation and song generation system for implementing the method of any one of claims 1 to 9, characterised in that, The method comprises the following modules: A data acquisition module is configured to acquire vocal data, vocal text and evaluation indexes of different singers as a sample set, and pre-process the sample set; A feature extraction module is connected with the data acquisition module, configured to extract features from the pre-processed sample set to obtain four-character features and context features, wherein the four-character features comprise acoustic layer features, physiological layer features, recognition layer features and aesthetic layer features; A model training module is connected with the feature extraction module, configured to train a four-character mapping model based on the four-character features and the context features, wherein the four-character mapping model is configured to output an acoustic signal based on the correlation between the four-character features; A music output module is connected with the model training module, configured to input to-be-processed vocal data of a target singer into the trained four-character mapping model, and output a target human voice.