Method and system for identifying style of popular music singer based on voiceprint recognition
By performing source separation and endpoint detection on popular music audio files, combined with linear predictive coding analysis and recursive feature decoupling, and mapping to a high-dimensional space for chaotic invariant estimation, the problem of insufficient voice separation accuracy in existing technologies is solved, and efficient and accurate identification of popular music singer styles is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUTURE DIMENSIONAL FILM CO LTD
- Filing Date
- 2026-03-30
- Publication Date
- 2026-05-19
AI Technical Summary
Existing techniques for identifying the style of pop singers suffer from insufficient precision in vocal separation, ineffective elimination of silent segments and transitional noise by endpoint detection, deviations in vocal tract parameter extraction, and failure to achieve precise separation and regularization of parameters during feature decoupling, resulting in low identification efficiency and poor accuracy.
By performing source separation and endpoint detection on popular music audio files, linear predictive coding analysis and recursive feature decoupling are carried out. The data is mapped to a high-dimensional space and chaotic invariant estimation is performed. A pre-set singer style database is used for similarity comparison to obtain dual-parameter chaotic features to identify style genres.
It achieves high accuracy and efficiency in identifying the style of pop music singers. By optimizing source separation and endpoint detection, it accurately extracts tract response parameters and stable tract cepstral feature sequences, thereby improving the accuracy and stability of the identification.
Smart Images

Figure CN122067526A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voiceprint identification technology, and in particular to a method and system for identifying the style of popular music singers based on voiceprint recognition. Background Technology
[0002] Existing techniques for identifying the style of pop singers lack sufficient precision in vocal separation, endpoint detection cannot effectively eliminate silent segments and transitional noise, there are deviations in the extraction of vocal tract parameters, and the feature decoupling process cannot achieve accurate separation and regularization of parameters.
[0003] Traditional identification methods do not perform high-dimensional phase space reconstruction and chaotic invariant calculation of vocal features, making it difficult to characterize the nonlinear dynamic characteristics of vocals. The style matching process is inefficient and has poor recognition accuracy. Therefore, how to improve the accuracy and efficiency of style identification for pop music singers has become an urgent problem to be solved. Summary of the Invention
[0004] This invention provides a method and system for identifying the style of popular music singers based on voiceprint recognition, in order to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides a method for identifying the style of popular music singers based on voiceprint recognition, comprising: The audio source of a pop music file is separated, and the isolated pure vocal track is segmented by endpoint detection to obtain continuous vocal segments of the pop music audio file. Linear predictive coding analysis is performed on continuous vocal segments to obtain the vocal tract response parameters of the continuous vocal segments; By recursively decoupling the vocal tract response parameters, a vocal tract cepstral feature sequence of continuous vocal segments is obtained. The vocal tract cepstral feature sequence is mapped to a high-dimensional space, and the chaotic invariant of the phase space trajectory obtained by the mapping is estimated to obtain the two-parameter chaotic feature of the phase space trajectory. Based on a pre-defined singer style database, similarity comparison is performed on dual-parameter chaotic features to identify the style and genre of singers in popular music audio files.
[0006] In a preferred embodiment, the sound source separation of popular music audio files includes: Perform a short-time Fourier transform on the audio file of popular music to obtain the time-spectrum energy distribution map of the audio file of popular music. Harmonic structure analysis was performed on the time-frequency energy distribution map to obtain the time-frequency distribution characteristics and repeating spectrum structure characteristics of the time-frequency energy distribution map; Based on the time-frequency distribution characteristics and repeating spectrum structure characteristics, the time-frequency energy distribution map is masked to obtain the vocal enhancement energy distribution map of popular music audio files; By combining the original phase information of the pop music audio file, an inverse short-time Fourier transform is performed on the vocal enhancement energy distribution map to obtain the pure vocal track of the pop music audio file.
[0007] In a preferred embodiment, the step of performing endpoint detection and segmentation on the separated pure vocal track to obtain continuous vocal segments of a pop music audio file includes: Waveform feature analysis was performed on the pure vocal track to obtain its energy envelope and zero-crossing rate; Based on the fluctuation characteristics of the energy envelope and the distribution characteristics of the zero-crossing rate, the activity range of the pure vocal track is determined to confirm the start and end boundaries of the vocal activity of the pure vocal track. Based on the start and end boundaries of human vocal activity, the pure vocal track is segmented, and the silent segments and transition noise in the segmented segments are removed to obtain continuous vocal segments of pop music audio files.
[0008] In a preferred embodiment, the step of performing linear predictive coding analysis on the continuous vocal segments to obtain the vocal tract response parameters of the continuous vocal segments includes: High-frequency attenuation compensation is applied to continuous sound segments to obtain component-enhanced sound segments of continuous sound segments; The component-enhanced speech segment is divided into stationary short-time analysis frames, and frame-weighted modulation is performed on the stationary short-time analysis frames to obtain the windowed speech frame sequence of the component-enhanced speech segment. Autocorrelation analysis was performed on the windowed speech frame sequence to construct the Toplitz matrix of the windowed speech frame sequence; By iterating through the Toplitz matrix, the linear prediction coefficient vector of the windowed speech frame sequence is obtained. By concatenating the linear prediction coefficient vectors into inter-frame parameters according to the time sequence, the tract response parameters of the continuous sound segments are obtained.
[0009] In a preferred embodiment, the recursive feature decoupling of the vocal tract response parameters to obtain the vocal tract cepstral feature sequence of continuous vocal segments includes: The recursive process of the linear prediction coefficient vector is traced back, and the prediction residual energy data generated during the linear prediction coding analysis is obtained. Based on the predicted residual energy data, the channel response parameters are subjected to energy equalization transformation to obtain the gain correction parameters of the channel response parameters. By weighted combination and recursion of the gain correction parameters, the cepstral characteristic sequence of the vocal tract of the continuous sound segment is obtained.
[0010] In a preferred embodiment, the recursive process of backtracking the linear prediction coefficient vector and obtaining the prediction residual energy data generated during the linear prediction coding analysis includes: By performing an inverse expansion of the Toplitz matrix, the reflection coefficients of each order of the Toplitz matrix can be obtained; Based on the reflection coefficients of each order, a lattice recursive analysis is performed on the windowed speech frame sequence to obtain the forward prediction error data of the windowed speech frame sequence; The amplitude of the forward prediction error data is accumulated to obtain the prediction residual energy data of the windowed speech frame sequence.
[0011] In a preferred embodiment, mapping the vocal tract cepstral feature sequence to a high-dimensional space includes: Autocorrelation decay analysis was performed on the vocal tract cepstral feature sequence. Based on the delay step size corresponding to the first decay to the initial value when the analyzed time-shift correlation increases with the delay step size, the delay time parameter of the vocal tract cepstral feature sequence was determined. Neighborhood evolution analysis was performed on the vocal tract cepstral feature sequence to obtain the embedding dimension parameter of the vocal tract cepstral feature sequence; Based on the delay time parameter and the embedding dimension parameter, the vocal tract cepstral feature sequence is embedded in a high-dimensional phase space to obtain the phase space trajectory of the vocal tract cepstral feature sequence.
[0012] In a preferred embodiment, the step of estimating the chaotic invariants of the mapped phase space trajectory to obtain the two-parameter chaotic characteristics of the phase space trajectory includes: Using the phase points in the phase space trajectory as reference points, neighborhood evolution tracking is performed on the reference points to construct the initial phase point pairs of the phase space trajectory; The phase point distance of the initial phase point pair is monitored to obtain the step size change sequence of the initial phase point pair; Based on the step size change sequence, a cumulative divergence rate convergence analysis is performed on the benchmark point, and the number of evolution steps corresponding to the first entry of the cumulative divergence rate into the preset stable threshold range is determined as the number of benchmark point tracking evolution steps. Sensitive dependency estimation is performed on the step size variation sequence, and the orbit divergence factor of the phase space trajectory is calculated. The formula for calculating the orbit divergence factor is as follows: ; In the formula, The orbital divergence factor of the phase space trajectory. This represents the total number of reference points selected from the phase space trajectory. In order to target the The number of evolutionary steps tracked by each reference point Delay time parameter of vocal tract cepstral feature sequence in phase space reconstruction For the first The initial time corresponding to each reference point For the first At the initial moment, each reference point and its nearest neighbor point... Spatial distance, For the process After the first step of evolution, the first A reference point and its nearest neighbor at time 1 Spatial distance, For summation operations; The phase space trajectory is divided into phase space grids, and the frequency distribution of the divided hypercube grids is analyzed to obtain the phase space complexity entropy of the phase space trajectory. Parametric coupling of the orbital divergence factor and the phase space complexity entropy yields the dual-parameter chaotic characteristics of the phase space trajectory.
[0013] In a preferred embodiment, the step of comparing the similarity of dual-parameter chaotic features based on a preset singer style database to identify the style and genre of singers in popular music audio files includes: Spatial distance analysis is performed on the dual-parameter chaotic features and the reference dual-parameter chaotic features in the preset singer style library to obtain the feature matching degree between the dual-parameter chaotic features and the reference dual-parameter chaotic features. Based on feature matching degree, similarity retrieval is performed on dual-parameter chaotic features to select the best candidate reference features for dual-parameter chaotic features. Based on the frequency statistics of candidate reference features, the style of singers in popular music audio files is determined, and the style and genre of the singers are obtained.
[0014] To address the aforementioned problems, the present invention also provides a pop music singer style identification system based on voiceprint recognition, the system comprising: The vocal segment extraction module is used to separate the sound sources of pop music audio files and perform endpoint detection and segmentation on the separated pure vocal tracks to obtain continuous vocal segments of pop music audio files. The vocal tract parameterization module is used to perform linear predictive coding analysis on continuous vocal segments to obtain the vocal tract response parameters of the continuous vocal segments; The cepstral feature decoupling module is used to perform recursive feature decoupling on the vocal tract response parameters to obtain the vocal tract cepstral feature sequence of continuous vocal segments. The chaotic feature extraction module is used to map the vocal tract cepstral feature sequence to a high-dimensional space and estimate the chaotic invariants of the mapped phase space trajectory to obtain the dual-parameter chaotic features of the phase space trajectory. The style matching and recognition module is used to compare the similarity of dual-parameter chaotic features based on a preset singer style library in order to identify the style and genre of singers in popular music audio files.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. By optimizing the sound source separation and endpoint detection process, this invention can obtain noise-free, pure, continuous sound segments. Through linear predictive coding analysis and recursive feature decoupling, it can accurately extract vocal tract response parameters and stable vocal tract cepstral feature sequences, thus solidifying the feature foundation for style identification.
[0016] 2. This invention maps cepstral features to a high-dimensional space and generates dual-parameter chaotic features. By comparing the similarity with a style library, it achieves rapid style determination, effectively improving the efficiency of style identification for pop music singers, while ensuring the accuracy and stability of the identification results. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a method for identifying the style of popular music singers based on voiceprint recognition, according to an embodiment of the present invention. Figure 2 A functional module diagram of a pop music singer style identification system based on voiceprint recognition provided in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a method for identifying the style of pop music singers based on voiceprint recognition. The execution subject of this method includes, but is not limited to, at least one electronic device that can be configured to execute the method provided in this application, such as a server or a terminal. In other words, the method for identifying the style of pop music singers based on voiceprint recognition can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cluster of cloud servers. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0020] Reference Figure 1The diagram shown is a flowchart illustrating a method for identifying the style of a pop music singer based on voiceprint recognition, according to an embodiment of the present invention. In this embodiment, the method for identifying the style of a pop music singer based on voiceprint recognition includes: The audio source of a pop music file is separated, and the isolated pure vocal track is segmented by endpoint detection to obtain continuous vocal segments of the pop music audio file. In this embodiment of the invention, the sound source separation of popular music audio files includes: Perform a short-time Fourier transform on the audio file of popular music to obtain the time-spectrum energy distribution map of the audio file of popular music. Harmonic structure analysis was performed on the time-frequency energy distribution map to obtain the time-frequency distribution characteristics and repeating spectrum structure characteristics of the time-frequency energy distribution map; Based on the time-frequency distribution characteristics and repeating spectrum structure characteristics, the time-frequency energy distribution map is masked to obtain the vocal enhancement energy distribution map of popular music audio files; By combining the original phase information of the pop music audio file, an inverse short-time Fourier transform is performed on the vocal enhancement energy distribution map to obtain the pure vocal track of the pop music audio file.
[0021] The process of segmenting the separated pure vocal track using endpoint detection to obtain continuous vocal segments from a pop music audio file includes: Waveform feature analysis was performed on the pure vocal track to obtain its energy envelope and zero-crossing rate; Based on the fluctuation characteristics of the energy envelope and the distribution characteristics of the zero-crossing rate, the activity range of the pure vocal track is determined to confirm the start and end boundaries of the vocal activity of the pure vocal track. Based on the start and end boundaries of human vocal activity, the pure vocal track is segmented, and the silent segments and transition noise in the segmented segments are removed to obtain continuous vocal segments of pop music audio files.
[0022] The popular music audio file is divided into continuous audio frames of fixed duration. A short-time Fourier transform operation is performed on each frame of audio signal to convert the time-domain audio signal of each frame into the corresponding frequency-domain signal. The signal energy values corresponding to different frequency points in the frequency-domain signal of each frame are counted. The frequency and energy data of all audio frames are fully integrated in time order and frequency dimension to generate the time-spectrum energy distribution map of the popular music audio file.
[0023] By traversing all frequency points and energy points in the time-spectrum energy distribution map, continuously tracking the distribution position of harmonic components at each frequency point, accurately determining the energy distribution pattern of different frequency bands, and fully recording the position and morphological characteristics of the periodic repetition of energy values, the time-frequency distribution characteristics and repetitive spectrum structure characteristics of the time-spectrum energy distribution map are extracted through systematic summarization and organization of the harmonic distribution law and energy repetition structure characteristics.
[0024] The time-frequency distribution features are matched one by one with the standard time-frequency features of the human voice signal, and the repeating spectral structure features are matched one by one with the standard spectral structure features of the human voice signal. The time-frequency regions in the time-frequency energy distribution map that are completely consistent with the matching results are accurately located. The energy preservation operation is performed on the successfully matched time-frequency regions, and the energy masking operation is performed on the unmatched time-frequency regions in the time-frequency energy distribution map. After the time-frequency masking operation is completed, the enhanced energy distribution map of the human voice of the pop music audio file is obtained.
[0025] The process involves extracting all phase information from the original time domain of a pop music audio file, precisely fusing the frequency domain energy data of the enhanced vocal energy distribution map with the original phase information frame by frame, performing an inverse short-time Fourier transform on all the fused frequency domain data, restoring each frame of frequency domain data to its corresponding time domain audio signal, and then stitching all the restored time domain audio frames together in their original time order to obtain the clean vocal track of the pop music audio file.
[0026] The waveform signal amplitude of the pure vocal track is accurately acquired frame by frame. The waveform amplitude of each frame is smoothed and stabilized, and the average energy value of the waveform in each frame is accurately calculated. All average energy values are continuously connected in chronological order to form a smooth energy change curve, thus obtaining the energy envelope of the pure vocal track. The total number of times the waveform signal in the pure vocal track crosses from positive to negative level and from negative to positive level is completely counted. The total number of times is accurately converted with the total duration of the pure vocal track to obtain the zero-crossing rate of the pure vocal track.
[0027] A fixed threshold for judging human voice energy is set based on the overall numerical distribution of the energy envelope. This threshold serves as the sole standard value for distinguishing between effective and ineffective human voice energy. Similarly, a fixed threshold for judging human voice zero-crossing rate is set based on the overall numerical distribution of the zero-crossing rate. This threshold serves as the sole standard value for distinguishing between human voice signals and non-human voice signals. All time periods of the pure human voice track are comprehensively traversed to identify consecutive time periods where both the energy envelope value and the zero-crossing rate value reach the human voice zero-crossing rate threshold. The starting point of these consecutive time periods is defined as the starting boundary of human voice activity, and the ending point is defined as the ending boundary of human voice activity. Ultimately, the starting and ending boundaries of human voice activity in the pure human voice track are determined.
[0028] The clean vocal track is cut without deviation according to the precisely defined start and end boundaries of vocal activity to obtain the corresponding initial vocal segments. A fixed silence energy threshold is set, which serves as the sole standard value to distinguish between silent and effective sound states. The portion of the initial vocal segment with an energy envelope value below the silence energy threshold is directly identified as a silent segment. The portion of the initial vocal segment with an energy envelope value above the silence energy threshold but below the vocal energy threshold is directly identified as transition noise. The identified silent segments and transition noise are completely removed from the initial vocal segments, and the remaining effective vocal parts are spliced together in their original time sequence to obtain the continuous vocal segments of the pop music audio file.
[0029] The beneficial effects are that the precise separation of vocals and accompaniment is achieved through layered and step-by-step time-domain and frequency-domain conversion and restoration operations; the purity of vocal signals is greatly improved by combining harmonic analysis and time-frequency masking; the accurate positioning of vocal intervals is achieved by relying on the precise determination of energy envelope and zero-crossing rate; invalid noise segments are completely eliminated; and high-quality continuous vocal segments are obtained, laying a solid data foundation for subsequent feature extraction and style identification.
[0030] Linear predictive coding analysis is performed on continuous vocal segments to obtain the vocal tract response parameters of the continuous vocal segments; In this embodiment of the invention, the step of performing linear predictive coding analysis on continuous vocal segments to obtain the vocal tract response parameters of the continuous vocal segments includes: High-frequency attenuation compensation is applied to continuous sound segments to obtain component-enhanced sound segments of continuous sound segments; The component-enhanced speech segment is divided into stationary short-time analysis frames, and frame-weighted modulation is performed on the stationary short-time analysis frames to obtain the windowed speech frame sequence of the component-enhanced speech segment. Autocorrelation analysis was performed on the windowed speech frame sequence to construct the Toplitz matrix of the windowed speech frame sequence; By iterating through the Toplitz matrix, the linear prediction coefficient vector of the windowed speech frame sequence is obtained. By concatenating the linear prediction coefficient vectors into inter-frame parameters according to the time sequence, the tract response parameters of the continuous sound segments are obtained.
[0031] By performing frequency band signal detection on a continuous sound segment, the high-frequency signal segment in the segment whose amplitude is reduced due to acquisition or transmission is located. Based on the standard compensation benchmark for high-frequency human voice signals, a preset fixed compensation amplitude is determined. The amplitude of each signal point in the high-frequency signal segment is synchronously increased according to the fixed compensation amplitude. Throughout the entire process of the increase operation, the amplitude of the mid- and low-frequency signals is kept in its original state without any change, so that the full-frequency signal amplitude of the continuous sound segment is restored to a balanced state. This completely solves the problem of signal feature loss caused by high-frequency signal attenuation. After completing the high-frequency attenuation compensation processing, the component-enhanced sound segment of the continuous sound segment is obtained.
[0032] Based on the stable characteristics of human voice signals, a preset fixed duration is set. The component-enhanced speech segments are then segmented without overlap according to this fixed duration. During the segmentation process, abnormal frames with drastic fluctuations in signal amplitude are filtered out, while frames with stable signal amplitude without abrupt changes are retained, forming multiple short-time analysis frames with stable signal states. For each stable short-time analysis frame, a preset weighting rule is used to perform amplitude adjustment. This rule keeps the signal amplitude at the center of the frame stable and gradually and smoothly reduces the amplitude towards the beginning and end edges of the frame, eliminating amplitude abrupt changes at the signal junctions between frames. After completing the weighted modulation of all frames, the processed frames are arranged strictly according to the original time order to obtain the windowed speech frame sequence of the component-enhanced speech segments.
[0033] Autocorrelation analysis was performed on each independent signal in the windowed speech frame sequence. The correlation between the signal and the original signal was calculated sequentially at different time offsets. Multiple sets of signal correlation values were obtained for each frame. The Toplitz matrix has the structural feature that the diagonal and parallel diagonal elements are equal. The calculated signal correlation values were accurately filled into the corresponding row and column positions of the matrix according to this symmetrical arrangement rule to ensure that the values of the corresponding positions in the horizontal and vertical directions of the matrix are completely consistent. After all the values were filled, the Toplitz matrix of the windowed speech frame sequence was successfully constructed.
[0034] Starting from the first-order initial order of linear predictive analysis, a stepwise recursive operation is performed on the Toplitz matrix. Each recursive operation is based on the result of the previous order operation, updating the feature values and related parameters inside the matrix. The recursive operation continues until the preset termination order of human voice linear prediction is reached. After the entire recursive process is completed, the final stable feature values corresponding to each frame of signal are extracted, and all feature values of the same frame are integrated and encapsulated to form the linear prediction coefficient vector corresponding to each frame in the windowed speech frame sequence.
[0035] By strictly adhering to the original temporal order of the windowed speech frame sequence, the linear prediction coefficient vectors corresponding to each frame in the sequence are seamlessly connected and combined in sequence to ensure that the end value of the coefficient vector of the previous frame is precisely matched with the beginning value of the coefficient vector of the next frame. There are no parameter interruptions, no numerical misalignments, and no order reversals throughout the process. After the complete splicing and integration of all frame linear prediction coefficient vectors, complete and continuous tract response parameters of the continuous speech segment are obtained.
[0036] The beneficial effects are that high-frequency attenuation compensation ensures the integrity of high-frequency components in the vocal segment without any loss, stable framing and weighted modulation improve the stability and coherence of signal analysis, autocorrelation analysis accurately constructs a standard Toplitz matrix, iteratively extracts linear prediction coefficient vectors, and orderly splices them to form complete and continuous duct response parameters, providing accurate and stable core parameter support for subsequent duct cepstral feature decoupling.
[0037] By recursively decoupling the vocal tract response parameters, a vocal tract cepstral feature sequence of continuous vocal segments is obtained. In this embodiment of the invention, the step of recursively decoupling the vocal tract response parameters to obtain the vocal tract cepstral feature sequence of continuous vocal segments includes: The recursive process of the linear prediction coefficient vector is traced back, and the prediction residual energy data generated during the linear prediction coding analysis is obtained. Based on the predicted residual energy data, the channel response parameters are subjected to energy equalization transformation to obtain the gain correction parameters of the channel response parameters. By weighted combination and recursion of the gain correction parameters, the cepstral characteristic sequence of the vocal tract of the continuous sound segment is obtained.
[0038] The recursive process of backtracking the linear prediction coefficient vector and obtaining the prediction residual energy data generated during the linear prediction coding analysis includes: By performing an inverse expansion of the Toplitz matrix, the reflection coefficients of each order of the Toplitz matrix can be obtained; Based on the reflection coefficients of each order, a lattice recursive analysis is performed on the windowed speech frame sequence to obtain the forward prediction error data of the windowed speech frame sequence; The amplitude of the forward prediction error data is accumulated to obtain the prediction residual energy data of the windowed speech frame sequence.
[0039] Following the complete process of gradually increasing order in the linear prediction coefficient vector generation, a comprehensive reverse decomposition and expansion operation is carried out on the Toplitz matrix. Starting from the highest termination order determined by the linear prediction analysis, the reverse derivation is performed step by step towards the initial first order. The internal structure of the matrix is independently decomposed for each order, and the characteristic correlation coefficients corresponding to the matrix at that order are extracted. The entire process ensures that there are no jumps in order and no omissions in coefficients, and finally, all the reflection coefficients of each order of the Toplitz matrix are obtained.
[0040] All extracted reflection coefficients of each order are used as the core calculation basis for lattice recursive analysis. A fully adapted lattice analysis structure is built based on the signal duration, amplitude distribution, and frequency band characteristics of the windowed speech frame sequence. The windowed speech frame sequence is processed frame by frame to perform continuous recursive calculation. Each calculation accurately compares the numerical difference between the original signal of the current frame and the predicted signal generated by the recursive model, and records each set of difference results completely. Finally, the forward prediction error data corresponding to all frames of the windowed speech frame sequence are extracted.
[0041] For each frame of independent forward prediction error data in the windowed speech frame sequence, the magnitude of the error signal is collected point by point. All the collected error amplitude values in the same frame are continuously accumulated and integrated. The accumulation process strictly follows the time order of the signal and does not miss any error amplitude point. After completing the accumulation operation of all amplitudes frame by frame, the complete and time-continuous prediction residual energy data of the windowed speech frame sequence is obtained.
[0042] Using the successfully acquired prediction residual energy data as a fixed energy calibration benchmark, a uniform adjustment operation is performed on the full-band energy distribution of the channel response parameters. The energy offset caused by signal acquisition, transmission, framing, weighting and other processes is corrected one by one, so that the energy values of each dimension of the channel response parameters are kept in a stable and balanced distribution. After completing all energy equalization transformation processing, the gain correction parameters corresponding to the channel response parameters are obtained.
[0043] According to a pre-set fixed weighting ratio adapted to human voice characteristics, the values of each dimension of the gain correction parameter are precisely matched and combined. Based on the parameters completed by the matching and combination, a step-by-step recursive integration operation is performed. The recursive process strictly follows the logical order of feature decoupling. The results generated by each step of the recursion are arranged in order according to the original time sequence of the continuous sound segments, and finally a complete and temporally coherent cepstral feature sequence of the continuous sound segments is obtained.
[0044] The beneficial effects are as follows: the reflection coefficients of each order are accurately extracted by complete inverse derivation of the Toplitz matrix; an adaptive lattice structure is built based on the reflection coefficients to complete the recursive analysis and reliably extract the forward prediction error data; the accurate prediction residual energy data is obtained by amplitude accumulation; the energy equalization transformation is completed based on the residual energy to generate stable gain correction parameters; the feature is accurately decoupled by weighted combination recursion to generate a regular and continuous tract cepstral feature sequence, which provides high-quality and stable feature support for subsequent high-dimensional phase space mapping.
[0045] The vocal tract cepstral feature sequence is mapped to a high-dimensional space, and the chaotic invariant of the phase space trajectory obtained by the mapping is estimated to obtain the two-parameter chaotic feature of the phase space trajectory. In this embodiment of the invention, mapping the vocal tract cepstral feature sequence to a high-dimensional space includes: Autocorrelation decay analysis was performed on the vocal tract cepstral feature sequence. Based on the delay step size corresponding to the first decay to the initial value when the analyzed time-shift correlation increases with the delay step size, the delay time parameter of the vocal tract cepstral feature sequence was determined. Neighborhood evolution analysis was performed on the vocal tract cepstral feature sequence to obtain the embedding dimension parameter of the vocal tract cepstral feature sequence; Based on the delay time parameter and the embedding dimension parameter, the vocal tract cepstral feature sequence is embedded in a high-dimensional phase space to obtain the phase space trajectory of the vocal tract cepstral feature sequence.
[0046] The process of estimating chaotic invariants from the mapped phase space trajectory to obtain the two-parameter chaotic characteristics of the phase space trajectory includes: Using the phase points in the phase space trajectory as reference points, neighborhood evolution tracking is performed on the reference points to construct the initial phase point pairs of the phase space trajectory; The phase point distance of the initial phase point pair is monitored to obtain the step size change sequence of the initial phase point pair; Based on the step size change sequence, a cumulative divergence rate convergence analysis is performed on the benchmark point, and the number of evolution steps corresponding to the first entry of the cumulative divergence rate into the preset stable threshold range is determined as the number of benchmark point tracking evolution steps. Sensitive dependency estimation is performed on the step size variation sequence, and the orbit divergence factor of the phase space trajectory is calculated. The formula for calculating the orbit divergence factor is as follows: ; In the formula, The orbital divergence factor of the phase space trajectory. This represents the total number of reference points selected from the phase space trajectory. In order to target the The number of evolutionary steps tracked by each reference point Delay time parameter of vocal tract cepstral feature sequence in phase space reconstruction For the first The initial time corresponding to each reference point For the first At the initial moment, each reference point and its nearest neighbor point... Spatial distance, For the process After the first step of evolution, the first A reference point and its nearest neighbor at time 1 Spatial distance, For summation operations; The phase space trajectory is divided into phase space grids, and the frequency distribution of the divided hypercube grids is analyzed to obtain the phase space complexity entropy of the phase space trajectory. Parametric coupling of the orbital divergence factor and the phase space complexity entropy yields the dual-parameter chaotic characteristics of the phase space trajectory.
[0047] Autocorrelation decay analysis was performed on the vocal tract cepstral feature sequence. The delay interval was gradually increased with a fixed step size of 0.01 seconds. The time-shift correlation between the sequence and the original sequence was calculated successively at different delay intervals. The correlation when the delay interval was zero was the initial correlation value. When the time-shift correlation first decayed to 0.368 times the initial correlation value, the current delay step size was determined as the delay time parameter of the vocal tract cepstral feature sequence.
[0048] Neighborhood evolution analysis is performed on the vocal tract cepstral feature sequence. Starting from dimension one, the spatial dimension value is gradually increased. For each dimension increase, the overlap ratio of the neighborhood phase points after evolution in that dimension is statistically analyzed to obtain the evolution stability. When the evolution stability reaches the preset fixed stability threshold of 0.9, the current spatial dimension value is determined as the embedding dimension parameter of the vocal tract cepstral feature sequence.
[0049] Based on the determined delay time parameters and embedding dimension parameters, a high-dimensional phase space embedding operation is performed on the vocal tract cepstral feature sequence. Each feature point in the sequence is offset and recombined according to the delay time parameters. An Euclidean phase space of the same dimension is constructed according to the embedding dimension parameters. All offset and recombined feature points are mapped one by one into this phase space and connected in time order to form a continuous phase point distribution trajectory, thus obtaining the phase space trajectory of the vocal tract cepstral feature sequence.
[0050] In the phase space trajectory, all phase points are evenly divided into intervals according to the time axis. One phase point is selected as the reference point in each interval. For each reference point, all phase points in the neighborhood are traversed and the spatial interval distance is calculated. The phase point with the smallest distance is selected as the nearest neighbor phase point. The reference point and the nearest neighbor phase point are fixedly combined to construct the initial phase point pair of the phase space trajectory.
[0051] For each initial phase pair, a continuous phase distance monitoring operation is performed. The spatial position change of the phase pair is tracked with a fixed evolution step size. After each evolution step is completed, the spatial distance value between the phase pairs is measured and recorded. The recording continues until the preset total number of two hundred evolution steps. All recorded distance values are arranged in the order of evolution to obtain the step size change sequence of the initial phase pair.
[0052] For each set of reference points in the step size change sequence, the number of evolutionary steps involved in the calculation is increased successively starting from the first evolution. For each additional evolutionary step, a cumulative divergence rate value is calculated using all the distance data from the reference point from the first step to the current step. Specifically, the natural logarithm of the ratio of the spatial distance between phase points after each evolution step to the initial spatial distance is obtained. This natural logarithm value is multiplied by the corresponding evolutionary step to obtain a single-step weighted value. All single-step weighted values within the current step size range are summed to obtain the numerator summed value. Then, the values of all evolutionary steps within the current step size range are squared and summed to obtain the denominator summed value. The numerator summed value is divided by the latter summed value to obtain the cumulative divergence rate value corresponding to the current evolutionary step size.
[0053] For each reference point, the cumulative divergence rate values calculated by successively increasing the evolution step size from the first step are arranged in ascending order of evolution step size to form a cumulative divergence rate sequence. Starting from the first value in the sequence, the sequence is traversed backward, and the absolute value of the difference between the current cumulative divergence rate value and the next cumulative divergence rate value is calculated. When the absolute value of the difference is less than 0.001 for five consecutive times, it is determined that the cumulative divergence rate has entered the stable threshold range, and the corresponding evolution step size is determined as the tracking evolution step number of the reference point.
[0054] Substitute the tracking evolution steps determined for each set of reference points in the above manner into the calculation process of the orbital divergence factor. First, count the total number of selected reference points, and set the corresponding tracking evolution steps for each reference point. Use the delay time parameter as the evolution time interval. First, calculate the ratio of the spatial distance between the phase points after each evolution step to the initial spatial distance and obtain the natural logarithmic value. Multiply the logarithmic value by the corresponding evolution step length to obtain the single-step weighted value. Accumulate all the single-step weighted values of the reference point from the first step to its tracking evolution step range to obtain the numerator accumulated value. Then, square the values of all evolution step lengths of the reference point from the first step to its tracking evolution step range and accumulate them to obtain the denominator accumulated value. Divide the numerator accumulated value by the latter accumulated result to obtain the divergence value of the reference point. Add the divergence values of all reference points and divide by the total number of reference points to obtain the arithmetic mean, and obtain the orbital divergence factor of the phase space trajectory.
[0055] A phase space grid partitioning operation is performed on the phase space trajectory. The phase space is divided into hypercube grids with a side length of 0.01 according to the embedding dimension parameter. The number of phase points in each hypercube grid is counted one by one. The proportion of the number of phase points falling into the grid to the total number of phase points in the phase space trajectory is used as the phase point distribution probability of the grid. The phase point distribution probability of each grid is multiplied by its natural logarithm, summed, and then the negative value is taken to obtain the information entropy value. This information entropy value is the phase space complexity entropy of the phase space trajectory.
[0056] The calculated orbit divergence factor and the phase space complex entropy are numerically bound according to a fixed feature combination rule. The two feature values correspond one-to-one to form a set of independent core feature data. This set of feature data is the dual-parameter chaotic feature of the phase space trajectory.
[0057] The beneficial effect is that, through the convergence analysis of the cumulative divergence rate, the number of tracking evolution steps for each reference point can be determined independently. This determination process is based on the stability criterion that the cumulative divergence rate value has five consecutive adjacent changes that are all less than 0.001. This ensures that the number of evolution steps for each reference point is selected at the minimum step size position when the divergence rate enters the stable plateau region, eliminating the estimation bias that may be introduced by subjectively setting a fixed step size, and making the calculation results of the orbital divergence factor more accurate and stable.
[0058] Based on a pre-defined singer style database, similarity comparison is performed on dual-parameter chaotic features to identify the style and genre of singers in popular music audio files.
[0059] In this embodiment of the invention, the step of comparing the similarity of dual-parameter chaotic features based on a preset singer style database to identify the style and genre of singers in popular music audio files includes: Spatial distance analysis is performed on the dual-parameter chaotic features and the reference dual-parameter chaotic features in the preset singer style library to obtain the feature matching degree between the dual-parameter chaotic features and the reference dual-parameter chaotic features. Based on feature matching degree, similarity retrieval is performed on dual-parameter chaotic features to select the best candidate reference features for dual-parameter chaotic features. Based on the frequency statistics of candidate reference features, the style of singers in popular music audio files is determined, and the style and genre of the singers are obtained.
[0060] Iterate through all the reference dual-parameter chaotic features in the preset singer style library, place the dual-parameter chaotic feature to be identified and each set of reference dual-parameter chaotic features in the same two-dimensional feature space, calculate the spatial distance between the feature points of the two sets one by one, and convert the spatial distance into the corresponding value according to the preset 0 to 1 value conversion rule. The smaller the spatial distance, the higher the converted value. The converted value is the feature matching degree between the dual-parameter chaotic feature and the reference dual-parameter chaotic feature.
[0061] Set 0.85 as a fixed feature matching degree screening threshold. Compare the feature matching degree of all reference dual-parameter chaotic features with this screening threshold one by one. Select reference dual-parameter chaotic features with a feature matching degree greater than or equal to 0.85. Sort the selected features in descending order of feature matching degree and retain the top 10 groups of features as candidate reference features for dual-parameter chaotic features.
[0062] Each candidate reference feature is labeled with its corresponding style category in the preset singer style library. The occurrence frequency of each style category is independently accumulated and statistically analyzed to determine the final cumulative occurrence frequency of each style category. The style category with the highest cumulative occurrence frequency is taken as the final judgment result, thus obtaining the style category of the singer in the pop music audio file.
[0063] The beneficial effects are that the matching degree between the feature to be identified and the reference feature is accurately calculated through two-dimensional feature space distance analysis, high-confidence candidate reference features are obtained by filtering and sorting based on fixed thresholds, and style classification is completed through frequency statistics. The overall comparison process is rigorous and efficient, which can quickly and accurately identify the style of the singer, and greatly improve the efficiency and accuracy of style identification of pop music singers.
[0064] like Figure 2 The diagram shown is a functional module diagram of a pop music singer style identification system based on voiceprint recognition, provided in an embodiment of the present invention.
[0065] The pop music singer style identification system 100 based on voiceprint recognition described in this invention can be installed in an electronic device. Depending on the functions implemented, the pop music singer style identification system 100 based on voiceprint recognition may include a vocal segment extraction module 101, a vocal tract parameterization module 102, a cepstral feature decoupling module 103, a chaotic feature extraction module 104, and a style matching and identification module 105. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0066] In this embodiment, the functions of each module / unit are as follows: The vocal segment extraction module 101 is used to separate the sound source of pop music audio files and perform endpoint detection and segmentation on the separated pure vocal tracks to obtain continuous vocal segments of pop music audio files. The vocal tract parameterization module 102 is used to perform linear predictive coding analysis on continuous vocal segments to obtain the vocal tract response parameters of the continuous vocal segments. The cepstral feature decoupling module 103 is used to perform recursive feature decoupling on the vocal tract response parameters to obtain the vocal tract cepstral feature sequence of continuous vocal segments. The chaotic feature extraction module 104 is used to map the vocal tract cepstral feature sequence to a high-dimensional space and estimate the chaotic invariants of the mapped phase space trajectory to obtain the dual-parameter chaotic features of the phase space trajectory. The style matching and recognition module 105 is used to perform similarity comparison on dual-parameter chaotic features based on a preset singer style library in order to identify the style and genre of singers in popular music audio files.
[0067] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0068] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0069] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0070] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0071] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for identifying the style of popular music singers based on voiceprint recognition, characterized in that, The method includes: The audio source of a pop music file is separated, and the isolated pure vocal track is segmented by endpoint detection to obtain continuous vocal segments of the pop music audio file. Linear predictive coding analysis is performed on continuous vocal segments to obtain the vocal tract response parameters of the continuous vocal segments; By recursively decoupling the vocal tract response parameters, a vocal tract cepstral feature sequence of continuous vocal segments is obtained. The vocal tract cepstral feature sequence is mapped to a high-dimensional space, and the chaotic invariant of the phase space trajectory obtained by the mapping is estimated to obtain the two-parameter chaotic feature of the phase space trajectory. Based on a pre-defined singer style database, similarity comparison is performed on dual-parameter chaotic features to identify the style and genre of singers in popular music audio files.
2. The method for identifying the style of popular music singers based on voiceprint recognition as described in claim 1, characterized in that, The process of separating the audio source of popular music files includes: Perform a short-time Fourier transform on the audio file of popular music to obtain the time-spectrum energy distribution map of the audio file of popular music. Harmonic structure analysis was performed on the time-frequency energy distribution map to obtain the time-frequency distribution characteristics and repeating spectrum structure characteristics of the time-frequency energy distribution map; Based on the time-frequency distribution characteristics and repeating spectrum structure characteristics, the time-frequency energy distribution map is masked to obtain the vocal enhancement energy distribution map of popular music audio files; By combining the original phase information of the pop music audio file, an inverse short-time Fourier transform is performed on the vocal enhancement energy distribution map to obtain the pure vocal track of the pop music audio file.
3. The method for identifying the style of popular music singers based on voiceprint recognition as described in claim 1, characterized in that, The process of segmenting the separated pure vocal track using endpoint detection to obtain continuous vocal segments from a pop music audio file includes: Waveform feature analysis was performed on the pure vocal track to obtain its energy envelope and zero-crossing rate; Based on the fluctuation characteristics of the energy envelope and the distribution characteristics of the zero-crossing rate, the activity range of the pure vocal track is determined to confirm the start and end boundaries of the vocal activity of the pure vocal track. Based on the start and end boundaries of human vocal activity, the pure vocal track is segmented, and the silent segments and transition noise in the segmented segments are removed to obtain continuous vocal segments of pop music audio files.
4. The method for identifying the style of popular music singers based on voiceprint recognition as described in claim 1, characterized in that, The linear predictive coding analysis of the continuous vocal segments yields the vocal tract response parameters of the continuous vocal segments, including: High-frequency attenuation compensation is applied to continuous sound segments to obtain component-enhanced sound segments of continuous sound segments; The component-enhanced speech segment is divided into stationary short-time analysis frames, and frame-weighted modulation is performed on the stationary short-time analysis frames to obtain the windowed speech frame sequence of the component-enhanced speech segment. Autocorrelation analysis was performed on the windowed speech frame sequence to construct the Toplitz matrix of the windowed speech frame sequence; By iterating through the Toplitz matrix, the linear prediction coefficient vector of the windowed speech frame sequence is obtained. By concatenating the linear prediction coefficient vectors into inter-frame parameters according to the time sequence, the tract response parameters of the continuous sound segments are obtained.
5. The method for identifying the style of popular music singers based on voiceprint recognition as described in claim 4, characterized in that, The recursive feature decoupling of the vocal tract response parameters to obtain the vocal tract cepstral feature sequence of continuous vocal segments includes: The recursive process of the linear prediction coefficient vector is traced back, and the prediction residual energy data generated during the linear prediction coding analysis is obtained. Based on the predicted residual energy data, the channel response parameters are subjected to energy equalization transformation to obtain the gain correction parameters of the channel response parameters. By weighted combination and recursion of the gain correction parameters, the cepstral characteristic sequence of the vocal tract of the continuous sound segment is obtained.
6. The method for identifying the style of popular music singers based on voiceprint recognition as described in claim 5, characterized in that, The recursive process of backtracking the linear prediction coefficient vector and obtaining the prediction residual energy data generated during the linear prediction coding analysis includes: By performing an inverse expansion of the Toplitz matrix, the reflection coefficients of each order of the Toplitz matrix can be obtained; Based on the reflection coefficients of each order, a lattice recursive analysis is performed on the windowed speech frame sequence to obtain the forward prediction error data of the windowed speech frame sequence; The amplitude of the forward prediction error data is accumulated to obtain the prediction residual energy data of the windowed speech frame sequence.
7. The method for identifying the style of popular music singers based on voiceprint recognition as described in claim 1, characterized in that, The process of mapping the vocal tract cepstral feature sequence to a high-dimensional space includes: Autocorrelation decay analysis was performed on the vocal tract cepstral feature sequence. Based on the delay step size corresponding to the first decay to the initial value when the analyzed time-shift correlation increases with the delay step size, the delay time parameter of the vocal tract cepstral feature sequence was determined. Neighborhood evolution analysis was performed on the vocal tract cepstral feature sequence to obtain the embedding dimension parameter of the vocal tract cepstral feature sequence; Based on the delay time parameter and the embedding dimension parameter, the vocal tract cepstral feature sequence is embedded in a high-dimensional phase space to obtain the phase space trajectory of the vocal tract cepstral feature sequence.
8. The method for identifying the style of popular music singers based on voiceprint recognition as described in claim 1, characterized in that, The process of estimating chaotic invariants from the mapped phase space trajectory to obtain the two-parameter chaotic characteristics of the phase space trajectory includes: Using the phase points in the phase space trajectory as reference points, neighborhood evolution tracking is performed on the reference points to construct the initial phase point pairs of the phase space trajectory; The phase point distance of the initial phase point pair is monitored to obtain the step size change sequence of the initial phase point pair; Based on the step size change sequence, a cumulative divergence rate convergence analysis is performed on the benchmark point, and the number of evolution steps corresponding to the first entry of the cumulative divergence rate into the preset stable threshold range is determined as the number of benchmark point tracking evolution steps. Sensitive dependency estimation is performed on the step size variation sequence, and the orbit divergence factor of the phase space trajectory is calculated. The formula for calculating the orbit divergence factor is as follows: ; In the formula, The orbital divergence factor of the phase space trajectory. This represents the total number of reference points selected from the phase space trajectory. In order to target the The number of evolutionary steps tracked by each reference point. Delay time parameter of vocal tract cepstral feature sequence in phase space reconstruction For the first The initial time corresponding to each reference point For the first At the initial moment, each reference point and its nearest neighbor point... Spatial distance, For the process After the first step of evolution, the first A reference point and its nearest neighbor at time 1 Spatial distance, For summation operations; The phase space trajectory is divided into phase space grids, and the frequency distribution of the divided hypercube grids is analyzed to obtain the phase space complexity entropy of the phase space trajectory. By parametrically coupling the orbital divergence factor and the phase space complexity entropy, the dual-parameter chaotic characteristics of the phase space trajectory are obtained.
9. The method for identifying the style of popular music singers based on voiceprint recognition as described in claim 1, characterized in that, The method, based on a pre-defined singer style database, performs similarity comparisons on dual-parameter chaotic features to identify the style and genre of singers in popular music audio files, including: Spatial distance analysis is performed on the dual-parameter chaotic features and the reference dual-parameter chaotic features in the preset singer style library to obtain the feature matching degree between the dual-parameter chaotic features and the reference dual-parameter chaotic features. Based on feature matching degree, similarity retrieval is performed on dual-parameter chaotic features to select the best candidate reference features for dual-parameter chaotic features. Based on the frequency statistics of candidate reference features, the style of singers in popular music audio files is determined, and the style and genre of the singers are obtained.
10. A pop music singer style identification system based on voiceprint recognition, characterized in that, The system for implementing the method for identifying the style of popular music singers based on voiceprint recognition as described in claim 1 includes: The vocal segment extraction module is used to separate the sound sources of pop music audio files and perform endpoint detection and segmentation on the separated pure vocal tracks to obtain continuous vocal segments of pop music audio files. The vocal tract parameterization module is used to perform linear predictive coding analysis on continuous vocal segments to obtain the vocal tract response parameters of the continuous vocal segments; The cepstral feature decoupling module is used to perform recursive feature decoupling on the vocal tract response parameters to obtain the vocal tract cepstral feature sequence of continuous vocal segments. The chaotic feature extraction module is used to map the vocal tract cepstral feature sequence to a high-dimensional space and estimate the chaotic invariants of the mapped phase space trajectory to obtain the dual-parameter chaotic features of the phase space trajectory. The style matching and recognition module is used to compare the similarity of dual-parameter chaotic features based on a preset singer style library in order to identify the style and genre of singers in popular music audio files.