Personalized customized ai voiceprint cloning method, system and storage medium thereof

By extracting voiceprint features through microphone acquisition and Mel frequency cepstral coefficient algorithm, and combining signal conversion frequency determination and feature dimension optimization, the problem of the inability of existing voiceprint cloning technology to personalize is solved, generating more natural and more recognizable personalized voice.

CN120783725BActive Publication Date: 2025-11-11HUNAN BOJI LIFE TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511280784.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-11-11
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Existing voiceprint cloning technology lacks personalized customization capabilities and cannot effectively capture and reproduce a user's unique acoustic characteristics, resulting in stiff and monotonous cloned voices that cannot meet personalized needs.

Method used

By collecting voice signals through a microphone and extracting voiceprint features using the Mel frequency cepstral coefficient algorithm, combined with signal conversion frequency determination and feature dimension optimization, accurate capture and personalized customization of voiceprint features are achieved. Combined with speech speed and pitch adjustment, voice that is more in line with the user's pronunciation habits is generated.

Benefits of technology

It achieves accurate mapping of the user's core voiceprint characteristics, improves the auditory similarity between the cloned voice and the original voice, meets diverse personalized needs, and makes the cloned voice more natural and more recognizable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783725B_ABST
    Figure CN120783725B_ABST
Patent Text Reader

Abstract

This invention discloses a personalized AI voiceprint cloning method, system, and storage medium, relating to the fields of artificial intelligence and speech processing technology. The personalized AI voiceprint cloning method includes the following steps: valid determination of speech signal conversion; voiceprint feature data acquisition and processing; and personalized voiceprint feature matching. This invention converts the collected speech signal into digital speech data, performs validity determination during the signal conversion process, then extracts voiceprint features from the qualified digital speech data using the Mel-frequency cepstral coefficient algorithm to obtain voiceprint feature data and perform voiceprint feature processing. Finally, it performs voiceprint feature conversion and combines it with a set speech text, achieving valid acquisition of speech data, accurate quantification of voiceprint feature extraction, and improved voiceprint feature recognition when combined with speech text. This effectively solves the problem of low voiceprint feature restoration in the existing voiceprint cloning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and speech processing technology, and in particular to a personalized AI voiceprint cloning method, system and storage medium. Background Technology

[0002] Personalized voiceprint cloning focuses on capturing and restoring the details of the target voiceprint to meet the needs of specific scenarios. The essence of voiceprint cloning is to extract the unique acoustic features of the target sound. Personalized voiceprint cloning first collects the user's voice samples, which need to cover different speech speeds, tones, and pronunciation scenarios to cover comprehensive voiceprint features. Then, it uses Mel spectrum to simulate the human ear's perception of sound and capture frequency information to extract core features from the preprocessed audio. Then, it performs personalized fine-tuning, using the target voiceprint data to fine-tune the model to adapt it to specific voiceprint features. Finally, it inputs text to generate the corresponding voiceprint speech.

[0003] For example, Chinese invention patent CN117672182B discloses a sound cloning method and system based on artificial intelligence, which includes: regularizing the original text and converting it into several sentences and words to be converted in sequence; obtaining the pinyin of the words to be converted and marking the pinyin of each character to obtain the first mark; splitting the initials and finals of the pinyin of the characters and assigning the first mark of the pinyin of the characters to the finals; marking the initials of the pinyin of the characters; determining the phoneme information according to preset rules; then recombining the word groups and determining the pause time between the recombined word groups according to the user's speaking speed; finally converting the characters and corresponding phoneme information into acoustic features and converting the acoustic features into target waveforms; and completing the sound cloning according to the target waveforms.

[0004] For example, Chinese Invention Patent No. CN115240630B discloses a method and system for converting Chinese text to personalized speech, which includes: extracting a fixed-length speaker feature embedding vector from the speaker's reference speech using a trained speaker encoder as the speaker's acoustic features; using a multi-speaker speech synthesis model Syn to convert the text to be converted into a Mel spectrogram corresponding to the speaker feature embedding vector; converting the Mel spectrum into the corresponding time-domain speech waveform, and outputting the final audio.

[0005] In existing technologies, most voiceprint cloning techniques can only simply copy the original speech and lack personalized customization of voiceprint style. Therefore, during the speech acquisition process, the single speech recognition algorithm in existing technologies ignores the uniqueness of individual voiceprints, making it impossible to effectively identify deep features that reflect the uniqueness of individual voiceprints, such as formant distribution and fundamental frequency changes. This results in the cloned speech signal appearing stiff and monotonous, failing to meet personalized needs, and consequently, failing to personalize the speech to meet specific requirements during the voiceprint cloning process. Summary of the Invention

[0006] To address the technical problem of existing technologies being unable to personalize voiceprints during the voiceprint cloning process, this invention provides a personalized AI voiceprint cloning method, system, and storage medium. The technical solution is as follows:

[0007] On the one hand, a personalized AI voiceprint cloning method is provided, which includes the following steps: acquiring the user's voice signal through a microphone, converting the acquired voice signal into digital voice data, and simultaneously determining the validity of the voice signal conversion process based on the signal conversion frequency to ensure the accuracy of subsequent voiceprint feature extraction and cloning. The voice signal carries the unique voiceprint features of the user, the phoneme sequence corresponding to the voice content, and natural prosody. The digital voice data is the input to the Mel-frequency cepstral coefficient algorithm, and the signal conversion frequency reflects the sampling accuracy level during the digital conversion process of the voice signal. The method then acquires the converted digital voice data, extracts voiceprint features using the Mel-frequency cepstral coefficient algorithm to obtain voiceprint feature data, and performs voiceprint feature processing. The Mel frequency cepstral coefficient algorithm is used to map the spectral features of digital speech data to the Mel frequency domain, which conforms to the characteristics of human hearing. Voiceprint features are used to distinguish the personalized features of different identified users. Voiceprint feature data represents a set of voiceprint features that have been quantized and can be recognized by computers. Voiceprint feature processing is used to denoise, standardize, and optimize the dimensionality of the acquired voiceprint feature data. Based on the results of voiceprint feature processing, voiceprint feature conversion is performed, and the converted voiceprint feature data is combined with the set speech text to complete the personalized AI voiceprint cloning. Voiceprint feature conversion is used to realize the mapping of voiceprint feature data with target voiceprint features that meet the personalized customization needs of the identified user. The set speech text represents the text content corresponding to the target voiceprint to be cloned, which is pre-specified by the identified user.

[0008] On the other hand, a personalized AI voiceprint cloning system is provided, comprising: a speech signal conversion validity determination module, a voiceprint feature data acquisition and processing module, and a personalized voiceprint feature matching module. The speech signal conversion validity determination module collects the user's voice signal through a microphone, converts the collected voice signal into digital speech data, and determines the validity of the voice signal conversion process based on the signal conversion frequency to ensure the accuracy of subsequent voiceprint feature extraction and cloning. The voiceprint feature data acquisition and processing module collects the converted digital speech data, extracts voiceprint features using the Mel-frequency cepstral coefficient algorithm, obtains voiceprint feature data, and performs voiceprint feature processing. The personalized voiceprint feature matching module performs voiceprint feature conversion based on the voiceprint feature processing results, and combines the converted voiceprint feature data with a set speech text to complete the personalized AI voiceprint cloning.

[0009] This application also provides a computer-readable storage medium storing a computer program, which is executed by a processor to provide a personalized AI voiceprint cloning method.

[0010] The beneficial effects of the technical solutions provided by the embodiments of the present invention include at least the following:

[0011] 1. The system collects the user's voice signal through a microphone, converts the collected voice signal into digital voice data, and determines the validity of the voice signal conversion process based on the signal conversion frequency to ensure the purity of the collected digital voice data. After the digital voice data is converted to qualified values, voiceprint features are extracted using the Mel-frequency cepstral coefficient algorithm to obtain voiceprint feature data. Voiceprint feature processing is then performed to effectively capture key features related to human auditory perception in the voice. Combined with voiceprint feature processing, more recognizable personalized voiceprint information can be extracted, reducing feature confusion. Based on the results of voiceprint feature processing, voiceprint feature conversion is performed, and the converted voiceprint feature data is combined with the set voice text to complete personalized AI voiceprint cloning. This achieves accurate mapping of the user's core voiceprint characteristics. The voice generated by combining the set text can highly restore the user's voice characteristics, making the cloned voice more similar to the original voice in terms of hearing.

[0012] 2. The validity of each dimension of the voiceprint feature data is analyzed through the signal-to-noise ratio (SNR) of the feature dimensions, with the minimum SNR range of the historical feature dimensions as the benchmark. For dimensions below the benchmark, the window increase is obtained by inputting the sliding window adjustment parameters into the optimization mapping set. If the ratio exceeds the limit, an alarm is triggered. The window is gradually increased step by step using the magnitude corresponding to the qualified window adjustment value as the adjustment step size for smooth processing until the SNR meets the standard. Dimensions that do not meet the standard are discarded, while those that meet the standard are marked as qualified feature dimensions and stored. For voices that cannot be personalized, the precise screening and retention of qualified feature dimensions can extract the unique voice characteristics of different users, breaking through the limitations of a single algorithm and capturing deep and unique features such as formants and fundamental frequencies. This avoids the stiffness and monotony of cloned voices, providing a rich and accurate feature foundation for personalized voice generation, making customized voices more in line with the user's pronunciation habits and style, and meeting diverse personalized needs.

[0013] 3. Speech Speed ​​Adjustment Judgment Process: Based on the matching process between syllable duration in the voiceprint feature data and the target text, the syllable speech speed is obtained. Combined with the average speech speed obtained during the voiceprint feature conversion period, a stepped threshold system of first, second, and third multiples is used to determine whether it is qualified and whether adjustment is needed. After adjustment, verification is performed until the ratio is within the first multiple range, and then the speech tone adjustment is judged again. This process significantly addresses the pain points of existing technologies: through accurate judgment and adjustment, it simulates the natural rhythm of the target voiceprint in different contexts, making up for the loss of details caused by noise in the original speech; the stepped threshold system combined with the verification mechanism avoids the limitations of a single algorithm, can reflect the uniqueness of individual voiceprints, makes cloned speech more natural, improves the voiceprint feature restoration accuracy, and meets personalized needs.

[0014] 4. Based on the voiceprint feature data after speech speed adjustment, obtain the fundamental frequency difference. If it is less than the set value, adjust the speech pitch; otherwise, complete the cloning to improve recognition accuracy. Speech pitch adjustment first obtains the energy value of each frequency point through frequency domain analysis, then obtains the total energy of each frequency band according to the frequency range, compares it with the set frequency band energy template, determines the energy ratio difference of each frequency band, and sends a calibration command to adjust the energy ratio of each frequency band accordingly. This process significantly addresses the pain points of existing technologies: by determining the fundamental frequency difference and adjusting the energy of each frequency band, it compensates for the loss of detail caused by noise in the original speech; it accurately adjusts the energy ratio of each frequency band, avoiding the limitations of a single algorithm, highlighting the uniqueness of individual voiceprints, making the cloned speech more natural, improving the voiceprint feature restoration accuracy, and meeting personalized needs. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart of a personalized AI voiceprint cloning method provided in an embodiment of the present invention;

[0017] Figure 2 A flowchart for determining the validity of speech signal conversion provided in an embodiment of the present invention;

[0018] Figure 3 This is a flowchart for determining the dimension of voiceprint features provided in an embodiment of the present invention;

[0019] Figure 4 A flowchart for determining speech speed adjustment provided in an embodiment of the present invention;

[0020] Figure 5 A flowchart for adjusting the energy ratio of each frequency band provided in an embodiment of the present invention;

[0021] Figure 6 This is a schematic diagram of the structure of a personalized AI voiceprint cloning system provided in an embodiment of the present invention. Detailed Implementation

[0022] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0023] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0024] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0025] like Figure 1The diagram shows a flowchart of a personalized AI voiceprint cloning method provided in an embodiment of the present invention. The method includes the following steps: acquiring the voice signal of an identified user through a microphone; converting the acquired voice signal into digital voice data; simultaneously determining the validity of the voice signal conversion process based on the signal conversion frequency to ensure the accuracy of subsequent voiceprint feature extraction and cloning; the voice signal carries the unique voiceprint features of the identified user, the phoneme sequence corresponding to the voice content, and natural prosody; the digital voice data is the input to the Mel-frequency cepstral coefficient algorithm; and the signal conversion frequency reflects the sampling accuracy level during the digital conversion of the voice signal; acquiring the converted digital voice data and extracting voiceprint features using the Mel-frequency cepstral coefficient algorithm to obtain voiceprint feature data; and performing voiceprint feature analysis. The process involves several steps: First, the Mel frequency cepstral coefficient algorithm maps the spectral features of digital speech data to the Mel frequency domain, which conforms to the characteristics of human hearing. Second, voiceprint features are used to distinguish the personalized characteristics of different identified users. Voiceprint feature data represents a quantized set of voiceprint features that can be recognized by a computer. Third, voiceprint feature processing is used to denoise, standardize, and optimize the dimensionality of the acquired voiceprint feature data. Finally, voiceprint feature transformation is performed based on the results of voiceprint feature processing. The transformed voiceprint feature data is then combined with a pre-defined voice text to complete a personalized AI voiceprint clone. Voiceprint feature transformation maps voiceprint feature data to target voiceprint features that meet the personalized customization needs of the identified user. The pre-defined voice text represents the text content corresponding to the target voiceprint to be cloned, as specified by the identified user.

[0026] In this embodiment, the reliability of the data foundation is ensured by accurately acquiring speech signals and strictly determining the effectiveness of conversion; the Mel-frequency cepstral coefficient algorithm is used to extract and optimize voiceprint features, enhancing the uniqueness and adaptability of the features; and feature conversion and cloning are completed by combining specified text, achieving highly personalized voiceprint generation that meets user needs. The beneficial effect of this implementation scheme is that it can provide users with personalized voiceprint cloning services, meet diverse needs in different scenarios, and improve the voice interaction experience; the working principle of this implementation scheme is to use voiceprint feature extraction algorithms and parameter adjustment technology to personalize voiceprint features, and then generate customized speech through speech synthesis technology.

[0027] The Mel-frequency cepstral coefficient algorithm is a voiceprint feature extraction method designed based on the characteristics of human hearing. Its core is to map the spectrum of the speech signal to the Mel-frequency domain. This process first performs a Fourier transform on the digital speech data to obtain the spectrum, then uses a Mel-filter bank to simulate the human ear's sensitivity to different frequencies, and finally obtains the Mel-frequency cepstral coefficients through a discrete cosine transform. Mel-frequency cepstral coefficients can effectively capture the spectral envelope features in speech, preserving personalized voiceprint information (such as pronunciation habits and vocal tract structure differences) while exhibiting good noise resistance. It is a key parameter characterizing the uniqueness of voiceprints, providing a precise feature foundation for subsequent personalized voiceprint cloning.

[0028] In one specific embodiment, such as Figure 2 The diagram shows a flowchart for determining the validity of voice signal conversion according to an embodiment of the present invention. The specific design logic is as follows: the matching range is located based on the signal conversion frequency. If it is in the valid range of 16kHz to 44.1kHz, it is determined to be qualified and the voiceprint features are extracted. If it is in the distortion range below 16kHz, it is determined to be distorted and a re-acquisition warning is triggered. If it is in the redundant range above 44.1kHz, a digital low-pass filter is used to filter out high-frequency components, and the signal conversion quality is re-acquired after verification. If it meets the standard, the first conversion optimization is completed and the availability is predicted. Otherwise, an acquisition warning is issued.

[0029] Further understanding is needed regarding the validity determination of speech signal conversion, which includes the following steps: Based on the signal conversion frequency acquired during the speech signal conversion process, it is compared with a set matching interval to locate the processing method corresponding to the interval. Specifically: If the acquired signal conversion frequency is within the first matching interval (usually set as a closed interval from 16kHz to 44.1kHz, i.e., the valid interval), the speech signal conversion process is deemed qualified, and voiceprint features are extracted using the Mel-frequency cepstral coefficient algorithm; if the acquired signal conversion frequency is within the second matching interval (usually set as an interval below 16kHz, i.e., the distortion interval), the speech signal conversion process is deemed distorted, and a re-acquisition warning is triggered to prompt the preset personnel to check the corresponding speech environment; if the acquired signal conversion frequency is within the third matching interval (usually set as an interval above 44.1kHz, i.e., the redundant interval), the speech signal conversion process is deemed redundant, and high-frequency components in the third matching interval are filtered out using a digital low-pass filter; the first matching interval, the second matching interval, and the third matching interval sequentially constitute the complete coverage range of the signal conversion frequency, and there is no overlap or gap between the intervals.

[0030] Based on the positioning results of the matching interval, the preprocessing optimization of the converted speech signal is realized. At the same time, a signal quality confirmation command is sent. The signal quality confirmation command is used to prompt the preset personnel to review the converted digital speech data in combination with the characteristics of the speech scene (such as the speech speed of daily conversation, the frequency range of professional pronunciation in specific fields, etc.) to confirm whether the current processing method accurately preserves the core speech features. This further ensures that while removing noise or redundancy, key acoustic information related to voiceprint recognition is retained, avoiding feature loss due to overprocessing. If the review confirms that the core speech features are accurately preserved, the first conversion optimization process is completed. Otherwise, a voice data acquisition warning is issued, such as prompting to check microphone sensitivity, change the acquisition environment (such as avoiding areas with strong electromagnetic interference), and adjust the user's speaking distance, to ensure the reliability of subsequent voiceprint processing data.

[0031] The process for confirming the accurate preservation of core speech features involves the following steps: Pre-selected personnel retrieve the converted digital speech data and compare it segment by segment with the speech scene characteristics. This includes checking whether the speech speed matches that of everyday conversations and whether the frequency range of professional pronunciations in specific fields is reasonable. Simultaneously, key acoustic information such as formants and fundamental frequencies are manually marked to check for loss during processing. If the personnel determine that the aforementioned core features are complete and consistent with the scene characteristics, then the data is confirmed as accurately preserved.

[0032] In this embodiment, by precisely defining three matching intervals—16kHz-44.1kHz (effective), <16kHz (distortion), and >44.1kHz (redundancy)—fine-grained control over conversion quality is achieved. The lower limit of 16kHz covers the core features of human speech (fundamental frequency and major harmonics are mostly distributed in this interval), preventing the loss of high-frequency features. The upper limit of 44.1kHz, referencing the Nyquist sampling theorem (satisfying twice the redundancy of the human ear's 20kHz hearing limit), ensures speech integrity while avoiding data redundancy. Combined with signal quality confirmation instructions guiding pre-selected personnel to review the speech based on contextual features, this approach achieves initial quality screening through technical means while incorporating human consideration of contextual characteristics. This effectively avoids the loss of key acoustic information due to over-processing, ensuring accurate preservation of core features related to voiceprint recognition during noise reduction or redundancy reduction.

[0033] Further, voiceprint features are extracted, including the following steps: The target speech signal is obtained from the corresponding digital speech data after the speech signal conversion process is qualified, and segmented into overlapping short frames with fixed durations (approximately 50% overlap between frames). The overlapping short frames are used to approximate the continuously dynamically changing acoustic features (such as frequency, amplitude, phase, etc.) in the target speech signal as short-time stationary signals. A windowing operation is performed based on the segmented overlapping short frames to obtain the short-frame speech signal corresponding to the overlapping short frames after the windowing operation, and a Fourier transform is performed. The Fourier transform is used to convert the time-domain signal in the short-frame speech signal into a frequency-domain signal to obtain the power spectrum of the corresponding short-frame speech signal. The windowing operation is used to reduce spectral leakage at both ends of the overlapping short frames.

[0034] The frequency axis of the power spectrum is introduced into the Mel frequency scale for filtering, and the Mel frequency cepstral coefficients are extracted to complete the initial extraction of voiceprint features. The voiceprint features after the initial extraction are obtained and feature filtering is performed to obtain voiceprint feature data. Feature filtering is used to remove redundant, noise-interfering, or low-correlation feature components (such as high-frequency interference features introduced by environmental noise, repetitive low-discrimination spectral information, etc.) from the voiceprint features, and retain the core features that play a key role in voiceprint recognition (such as fundamental frequency range, formant distribution, spectral envelope details, etc.).

[0035] In this embodiment, by segmenting overlapping short frames, detailed features such as frequency and amplitude that change over time in the speech signal are effectively preserved. Mel frequency cepstral coefficients are extracted to specifically capture key spectral features related to human speech, laying the foundation for the feature selection process by eliminating redundant feature components. The overall process, through segmented processing of the speech signal, spectrum conversion, and feature extraction, achieves accurate capture of the core voiceprint features of human voice, effectively improving the recognition of voiceprint features of different users, and providing key support for the accuracy of subsequent personalized voiceprint cloning, ensuring accurate capture of dynamic acoustic features.

[0036] Further, voiceprint feature processing includes the following steps: First, the acquired voiceprint feature data is analyzed for feature dimension validity using the feature dimension signal-to-noise ratio (SNR). The feature dimension SNR is used to quantify the difference between the effective signal and noise signal of each dimension feature in the voiceprint feature data, i.e., the ratio of the effective information energy contained in that dimension feature that reflects the uniqueness of the target user's voiceprint to the noise energy (such as ambient background noise) mixed in that dimension. Second, the minimum historical feature dimension SNR range is used as the benchmark for feature dimension validity analysis, and voiceprint feature dimension determination is performed. The voiceprint feature dimension determination is used to judge the validity of the corresponding feature dimension in the voiceprint features. The minimum historical feature dimension SNR range is represented by the sum and average of the minimum historical feature dimension SNR ranges in each historical voiceprint feature processing process in the database.

[0037] Specifically, the calculation method for the signal-to-noise ratio (SNR) of a feature dimension is first clarified. For feature dimensions with different units of measurement, such as fundamental frequency, formant frequency, and Mel-frequency cepstral coefficient, a unified calculation logic of relative difference quantification is adopted. That is, for a certain dimension to be analyzed, the overall difference of the feature of that dimension between different users is calculated by comparing the feature of the target user with a sufficient number of similar users (such as the dispersion of the feature mean among users). This quantifies the effective information energy of the uniqueness of the target user's voiceprint. The higher the degree of difference, the more sufficient the effective information energy. Then, the fluctuation of the feature of that dimension in the voice samples recorded by the target user in different scenarios and at different times is analyzed (such as the deviation of the feature value within the user from its own mean). This quantifies the noise energy mixed in the dimension. The lower the degree of fluctuation, the smaller the noise energy. Then, the effective information energy obtained by the above quantification is compared with the noise energy to obtain the feature dimension SNR of that dimension. This ensures that feature dimensions with different units of measurement can complete the SNR calculation through a unified logic.

[0038] Next, the method for determining the benchmark for feature dimension validity analysis is clarified: The minimum signal-to-noise ratio (SNR) range of each feature dimension analyzed in all past voiceprint feature processing cases is retrieved from the database. These minimum values ​​are summed and averaged; the resulting value is the benchmark for the current feature dimension validity analysis. Finally, voiceprint feature dimensions are determined: the calculated SNR of each feature dimension is compared with the benchmark. If the SNR of a dimension is greater than or equal to the benchmark, the dimension is considered to contain valid voiceprint information and is retained; if the SNR is less than the benchmark, the dimension is considered redundant (significantly affected by environmental interference or extraction errors) and is removed. This completes dimension optimization, ensuring that technicians in the relevant technical field can reproduce the operation based on this logic.

[0039] In this embodiment, the problem of calculating the signal-to-noise ratio of different dimensional features such as fundamental frequency and Mel-frequency cepstral coefficients is solved by unifying the calculation logic through relative difference quantification. This ensures that the validity analysis standards of each dimension are consistent and avoids judgment bias caused by differences in units. By quantifying the effective information energy with the feature differences between users and the noise energy with the feature fluctuations within the same user, the effective dimensions containing the uniqueness of the target user's voiceprint can be accurately screened, improving the purity of voiceprint feature data. The validity analysis benchmark is determined based on historical data, making the dimension judgment more objective and stable, reducing subjective errors. At the same time, the optimized feature dimensions reduce the computational load of subsequent voiceprint cloning processing, providing reliable feature support for high-quality voiceprint cloning.

[0040] like Figure 3The diagram shows a flowchart for determining the voiceprint feature dimension according to an embodiment of the present invention. The specific design logic is as follows: First, the feature dimension and preset parameters below the historical minimum signal-to-noise ratio are input into the optimization mapping set to obtain the increase of the sliding window for noise reduction; if the ratio exceeds the preset maximum threshold, an alarm is issued; otherwise, it is marked as a qualified value; the window is gradually increased according to the qualified value range for smooth processing until the signal-to-noise ratio is not less than the historical minimum value and then stops; if it still does not meet the standard after optimization, a rejection prompt is issued; otherwise, it is marked as a qualified feature dimension and stored in the database.

[0041] Further understanding is needed regarding the voiceprint feature dimension determination, which includes the following steps: inputting the dimensional features in the voiceprint feature data corresponding to the minimum value of the historical feature dimension signal-to-noise ratio range below the minimum value into the sliding window optimization mapping set used to improve the feature signal-to-noise ratio, obtaining the corresponding sliding window increase magnitude, so as to effectively remove environmental noise and accurately extract effective voiceprint features. The preset sliding window adjustment parameters are represented by the result of summing and averaging the historical sliding window adjustment parameters in the historical voiceprint feature processing process in the database.

[0042] It should be noted that the sliding window optimized mapping set is generated based on historical voiceprint processing data. Using voiceprint feature dimension data from different scenarios as samples, the signal-to-noise ratio (SNR) of the feature dimensions and historical adjustment parameters are taken as inputs, and the increase in the sliding window is taken as the output. The model is trained using machine learning algorithms (such as regression analysis and neural networks). During training, the model parameters are iteratively adjusted with the goal of optimizing the SNR, so that the mapping set can adaptively output the matching window increase based on the input features, forming a stable mapping relationship.

[0043] The algorithm determines whether the increase in the sliding window exceeds the maximum allowable value. If so, it sends an alarm indicating that the window adjustment exceeds the limit. Otherwise, it marks the increase in the sliding window as a qualified adjustment value. The maximum allowable value of the sliding window is represented by the average of the historical maximum allowable values ​​in each historical voiceprint feature processing process in the database. Based on the magnitude corresponding to the qualified adjustment value, the sliding window is gradually increased for smoothing. The algorithm continuously monitors the recalculated feature dimension signal-to-noise ratio. When the feature dimension signal-to-noise ratio is not less than the minimum historical feature dimension signal-to-noise ratio range, the algorithm stops execution. After the feature dimension signal-to-noise ratio is optimized, if the feature dimension signal-to-noise ratio is still less than the minimum historical feature dimension signal-to-noise ratio range, a feature removal prompt is sent. Otherwise, the feature dimension after the sliding window adjustment is marked as a qualified feature dimension and synchronized to the feature database for storage.

[0044] The sliding window increase magnitude, maximum allowable value of the sliding window, and adjustment step size are progressively constrained and executed, forming the core logic of sliding window optimization. The maximum allowable value of the sliding window is the boundary threshold for adjustment, determined by the average of the historical maximum allowable values. It is used to limit the extreme range of window increase and prevent the loss of voiceprint feature details (such as subtle tone changes) due to excessive smoothing caused by an overly large window. The sliding window increase magnitude is the initial adjustment requirement calculated based on the current signal-to-noise ratio gap in the feature dimension. It represents the theoretically required increase magnitude. It is compared with the maximum allowable value. If it does not exceed the limit, it becomes a qualified window adjustment value. If it exceeds the limit, an alarm is triggered to avoid invalid adjustment. The adjustment step size is the actual unit of window increase, directly determined by the qualified sliding window increase magnitude. In the smoothing process, the window is increased step by step according to this step size. The signal-to-noise ratio is recalculated after each adjustment until the target is met.

[0045] This closed-loop adjustment mechanism not only controls the adjustment bottom line by setting the maximum allowable value, but also clarifies the adjustment target by increasing the magnitude. Finally, it achieves precise and controllable window optimization through step size, ensuring a balance between noise reduction effect and feature integrity. At the same time, combined with the window adjustment over-limit alarm mechanism, it can effectively ensure the safety and rationality of the feature optimization process, fundamentally reducing the interference of invalid features on cloning accuracy, and further improving the accuracy of personalized AI voiceprint cloning.

[0046] In one specific embodiment, such as Figure 4 The diagram shows a flowchart of the speech speed adjustment determination provided in an embodiment of the present invention. The specific design logic is as follows: Based on the matching process between the syllable duration in the voiceprint feature data and the target text, the syllable speech speed is obtained. The average speech speed is obtained by combining the voiceprint feature conversion time period. If the syllable speech speed is within the range of 0.7-1.3 times the average speech speed, it is deemed qualified and no adjustment is required. If it is less than 0.7 times, it is deemed slow, and the time axis stretching ratio is increased. If it is greater than 1.3 times, it is deemed fast, and the stretching ratio is decreased. After adjustment, the two and the ratio are re-acquired. If they are within the range of 0.7-1.3 times, it is deemed qualified and the current ratio is retained. Otherwise, the corresponding adjustment operation is repeated until the ratio is within this range.

[0047] Further understanding is needed regarding voiceprint feature conversion, which includes the following steps: obtaining qualified feature dimensions for speech speed adjustment and speech pitch adjustment. Speech speed adjustment is used to simulate the natural expression rhythm of the target voiceprint in different contexts based on the syllable duration in the voiceprint feature data. Speech pitch adjustment is used to determine the recognizability of the voiceprint features based on the fundamental frequency information in the voiceprint feature data after speech speed adjustment.

[0048] The specific process for determining voice speed adjustment includes:

[0049] L1, based on syllable speech velocity and average speech velocity, makes a judgment: Syllable speech velocity is obtained by matching the syllable duration in the voiceprint feature data with the target text (the content to be synthesized). The matching process involves matching the syllables extracted from the target text with the syllable durations extracted from the voiceprint feature data. Simultaneously, combined with the segmented voiceprint feature transition time periods, the average speech velocity is obtained. Syllable duration represents the duration of a single syllable corresponding to the target voiceprint feature within the voiceprint feature transition time period, from the start to the end of pronunciation. Syllable speech velocity represents the speed of pronunciation of a single syllable corresponding to the target voiceprint feature within the voiceprint feature transition time period. Speech rate, which is the proportion of time corresponding to the pronunciation duration of each syllable, can be quantified as "syllables / second". It intuitively reflects the speed and rhythm of the target voiceprint when pronouncing specific syllables (e.g., if the pronunciation of a syllable takes 0.5 seconds, then its speech rate is 2 syllables / second). Average speech rate represents the average pronunciation rate of all syllables of the target voiceprint feature during the voiceprint feature transition period, that is, the ratio of the total number of syllables to the total duration during that period (quantified as "total number of syllables / total duration (seconds)"). It is used to reflect the overall speaking rhythm speed of the target voiceprint during a specific period (e.g., if a transition period contains 10 syllables and the total duration is 5 seconds, then the average speech rate is 2 syllables / second).

[0050] If the obtained syllable speech speed is within the first multiple of the average speech speed (usually 0.7-1.3 times), it is considered acceptable and no speech speed adjustment is needed. If the obtained syllable speech speed is lower than the second multiple of the average speech speed (usually 0.7 times), it is considered slow, and the time axis stretching ratio is increased to slow down the cloned speech speed. If the obtained syllable speech speed is higher than the third multiple of the average speech speed (usually 1.3 times), it is considered fast, and the time axis stretching ratio is decreased to speed up the cloned speech speed. The first, second, and third multiples constitute a stepped threshold system for speech speed determination, and the values ​​increase sequentially.

[0051] L2, Verification of speech speed adjustment effect: After speech speed adjustment, the syllable speech speed and average speech speed are reacquired, and the corresponding ratio is obtained. If the ratio is within the first multiple range, the speech speed adjustment is deemed qualified and the current time axis stretch ratio is retained. At the same time, the speech tone adjustment is judged. If the ratio is lower than the second multiple or higher than the third multiple, the corresponding stretch ratio adjustment operation is repeated until the ratio is within the first multiple range.

[0052] L3, Voice Pitch Adjustment Judgment: Based on the voiceprint feature data after voice speed adjustment, the fundamental frequency difference is obtained. If the fundamental frequency difference is less than the set fundamental frequency difference, the voice pitch is adjusted. Otherwise, the target voiceprint feature is obtained based on the result of voiceprint feature conversion, and a personalized AI voiceprint clone is completed to improve the recognition of voiceprint features. The fundamental frequency difference represents the difference between the maximum and minimum values ​​of the fundamental frequency of the voiceprint feature data. The set fundamental frequency difference is represented by the sum and average of the historical fundamental frequency differences in the historical voice speed adjustment process in the database.

[0053] In this embodiment, a tiered threshold system is used to achieve fine-grained control over speech speed. The acceptable range is 0.7-1.3 times the normal speed, referencing the fluctuation range of natural speech speed in everyday human conversation, preserving the rhythm of the speech while filtering out obvious abnormalities. 0.7 times and 1.3 times are used as the thresholds for judging excessively slow and excessively fast speech, respectively, forming clearly defined adjustment trigger conditions.

[0054] This example utilizes precise judgment in the L1 stage to perform targeted adjustments to the timeline stretching ratio; the L2 stage's effect verification mechanism ensures that the adjusted speech rate is within a reasonable range, avoiding over-correction; and the L3 stage combines fundamental frequency difference for pitch adjustment judgment, preventing pitch distortion caused by speech rate adjustments. The overall process balances the naturalness and standardization of the speech, significantly improving the fluency and recognizability of personalized voiceprint clones through dynamic adjustments and multi-dimensional verification, making the generated speech more closely match the target user's expression habits.

[0055] Further, the voice tone adjustment includes the following steps: performing frequency domain analysis based on the acquired original voiceprint feature data to obtain the energy value of each frequency point; obtaining the total energy of each frequency band according to the acquired frequency band range, which includes the first frequency band (high frequency band above 2000Hz), the second frequency band (mid frequency band of 300-2000Hz), and the third frequency band (low frequency band below 300Hz); comparing the total energy of each frequency band with the set frequency band energy template respectively, determining the energy ratio difference of each frequency band, and sending a frequency band energy ratio calibration command, which includes the set first frequency band (i.e., high frequency band 30%-40%), the second frequency band (i.e., mid frequency band 35%-50%), and the third frequency band (i.e., low frequency band 20%-25%).

[0056] In this embodiment, unlike the existing technology that directly adjusts the overall tone, the innovation of this solution lies in achieving precise control through frequency band matching: first, the speech is divided into three segments—high frequency, mid frequency, and low frequency—and then matched and calibrated with the energy templates of the corresponding frequency bands. This layered matching mechanism overcomes the limitation of traditional overall adjustment, which struggles to take into account the characteristics of each frequency band. It can specifically optimize the energy ratio of different frequency bands: it ensures clarity through high-frequency matching, restores the core timbre through mid-frequency matching, and maintains tone stability through low-frequency matching, enabling each frequency band to work synergistically. While preserving the uniqueness of the voiceprint, it achieves more natural and precise tone optimization.

[0057] like Figure 5 The diagram shows a flowchart of the energy ratio adjustment for each frequency band provided in an embodiment of the present invention. The specific design logic is as follows: the energy ratio is adjusted through three frequency bands. In the first frequency band, when the energy ratio difference is positive, the high-frequency energy output is reduced, and when it is negative, the high-frequency energy is increased to match the bright and penetrating tonal characteristics. In the second frequency band, when the difference is positive, the mid-frequency energy is reduced, and when it is negative, the mid-frequency energy is increased to ensure that the speech clarity and recognizability meet the target template. In the third frequency band, when the difference is positive, the low-frequency energy is reduced, and when it is negative, the low-frequency energy is increased to adjust the thickness and stability of the speech.

[0058] It is further important to understand that the energy ratio difference for each frequency band includes the energy ratio difference for the first frequency band, the energy ratio difference for the second frequency band, and the energy ratio difference for the third frequency band. The energy ratio difference for the first frequency band represents the difference between the energy proportion of the first frequency band within the frequency band range and the preset energy proportion within the first frequency band. The energy ratio difference for the second frequency band represents the difference between the energy proportion of the second frequency band within the frequency band range and the preset energy proportion within the second frequency band. The energy ratio difference for the third frequency band represents the difference between the energy proportion of the third frequency band within the frequency band range and the preset energy proportion within the third frequency band. The energy ratio is calculated by power spectral density to obtain the energy of the corresponding frequency band and then proportionally calculated with the total frequency band energy. The preset energy ratio is represented by the sum and average of the historical energy ratios from the historical speech tone adjustment process in the database. The frequency band energy ratio calibration command is used to adjust the energy ratio of each frequency band specifically based on the energy ratio differences of the first, second, and third frequency bands.

[0059] If the energy ratio difference of the first frequency band is positive (actual ratio higher than the preset 30%-40%), then reduce the energy output of the high frequency band; if it is negative (actual ratio insufficient), then increase the energy of the high frequency band to match the bright and penetrating tonal characteristics. If the energy ratio difference of the second frequency band is positive (actual ratio higher than the preset 35%-50%), then reduce the energy of the mid frequency band; if it is negative, then increase the energy of the mid frequency band to ensure that the speech clarity and recognizability meet the target template. If the energy ratio difference of the third frequency band is positive (actual ratio higher than the preset 20%-25%), then reduce the energy of the low frequency band; if it is negative, then increase the energy of the low frequency band to adjust the thickness and stability of the speech.

[0060] This example uses the frequency band energy ratio calibration command to achieve dynamic balance of energy in each frequency band, so that the adjusted voiceprint features match the set frequency band energy template in terms of pitch and intensity fluctuations, accurately restoring the personalized tone style of the target voiceprint.

[0061] In this embodiment, dynamic calibration is achieved by accurately responding to the frequency band energy ratio difference, which improves the fineness and personalized restoration capability of voice tone adjustment. The difference between the actual ratio and the preset range is adjusted in the positive or negative direction, which can accurately control the influence of each frequency band on the tone, ensuring that the adjusted voiceprint features fit the unique texture of the target voiceprint and avoid timbre distortion caused by adjustment imbalance. The frequency band energy ratio calibration command realizes the coordinated balance of each frequency band, so that the adjusted voiceprint is highly consistent with the set template in terms of pitch and intensity fluctuations, ensuring the natural and smooth tone features.

[0062] like Figure 6 The diagram shows the system structure of the personalized AI voiceprint cloning method provided in this embodiment of the invention. The system includes: a language signal conversion validity determination module, a voiceprint feature data acquisition and processing module, and a personalized voiceprint feature matching module. The language signal conversion validity determination module collects the user's voice signal through a microphone, converts the collected voice signal into digital voice data, and determines the validity of the voice signal conversion process based on the signal conversion frequency to ensure the accuracy of subsequent voiceprint feature extraction and cloning. The voiceprint feature data acquisition and processing module collects the converted digital voice data, extracts voiceprint features using the Mel frequency cepstral coefficient algorithm, obtains voiceprint feature data, and performs voiceprint feature processing. The personalized voiceprint feature matching module performs voiceprint feature conversion based on the results of voiceprint feature processing and combines the converted voiceprint feature data with a set voice text to complete the personalized AI voiceprint cloning.

[0063] In this embodiment, the output of the speech signal conversion efficiency evaluation module serves as the input benchmark for the voiceprint feature data acquisition and processing module. Its quality assessment provides a preliminary guarantee for the accuracy of subsequent feature extraction. Meanwhile, the optimized feature data from the voiceprint feature data acquisition and processing module provides a high-quality data source for the feature conversion and synthesis operations of the personalized voiceprint feature matching module, ensuring the accuracy and stability of the feature mapping process. This modular, tightly coupled architecture effectively avoids the cumulative impact of quality defects in preceding stages on subsequent processes. Furthermore, through the specialized processing mechanisms of each module, it significantly improves the overall voiceprint cloning accuracy of the system, ultimately achieving efficient and accurate operation of the personalized AI voiceprint cloning system from signal acquisition and feature processing to result output.

[0064] This application provides a computer-readable storage medium storing a computer program, which is executed by a processor to provide a personalized AI voiceprint cloning method.

[0065] It should be understood that the general-purpose processor can be a microprocessor or any conventional processor. The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuitry), firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0066] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0067] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0068] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0069] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0070] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0071] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0072] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0073] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0074] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0075] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A personalized AI voiceprint cloning method, characterized in that, Includes the following steps: S1. The voice signal of the identified user is collected through a microphone and converted into digital voice data. At the same time, the validity of the voice signal conversion process is determined based on the signal conversion frequency to ensure the accuracy of subsequent voiceprint feature extraction and cloning. The voice signal is used to carry the unique voiceprint features of the identified user, the phoneme sequence corresponding to the voice content, and the natural rhythm. The digital voice data is the input of the Mel frequency cepstral coefficient algorithm. The signal conversion frequency is used to reflect the sampling accuracy level in the digital conversion process of the voice signal. S2, Collect qualified digital speech data, extract voiceprint features using the Mel frequency cepstral coefficient algorithm to obtain voiceprint feature data, and perform voiceprint feature processing. The Mel frequency cepstral coefficient algorithm is used to map the spectral features of digital speech data to the Mel frequency domain that conforms to the hearing characteristics of the human ear. The voiceprint features are used to distinguish the personalized features of different identified users. The voiceprint feature data represents a set of voiceprint features that have been quantized and can be recognized by a computer. The voiceprint feature processing is used to denoise, standardize, and optimize the dimensions of the obtained voiceprint feature data. S3, based on the results of voiceprint feature processing, perform voiceprint feature conversion, and combine the converted voiceprint feature data with the set voice text to complete the personalized AI voiceprint cloning. The voiceprint feature conversion is used to realize the mapping of voiceprint feature data with target voiceprint features that meet the personalized customization needs of the user. The set voice text represents the text content corresponding to the target voiceprint to be cloned, which is pre-specified by the user. The voiceprint feature processing includes the following steps: The feature dimension signal-to-noise ratio is used to analyze the effectiveness of the acquired voiceprint feature data. The feature dimension signal-to-noise ratio is used to quantify the difference between the effective signal and the noise signal of each dimension feature in the voiceprint feature data. The minimum signal-to-noise ratio range of historical feature dimensions is used as the benchmark for feature dimension validity analysis, and voiceprint feature dimension determination is performed. The voiceprint feature dimension determination is used to judge the validity of the corresponding feature dimension in the voiceprint features. The determination of the voiceprint feature dimensions includes the following steps: The dimensional features in the voiceprint feature data corresponding to the minimum value of the signal-to-noise ratio of the feature dimension below the historical feature dimension are input into the sliding window optimization mapping set used to improve the feature signal-to-noise ratio, and the corresponding increase in sliding window value is obtained. Determine whether the increase in the sliding window exceeds the maximum allowable value of the sliding window. If so, send a window adjustment over-limit alarm; otherwise, mark the increase in the sliding window as a qualified window adjustment value. Based on the magnitude corresponding to the qualified window adjustment value as the adjustment step size, the sliding window is gradually increased for smoothing. The recalculated feature dimension signal-to-noise ratio is continuously monitored. When the feature dimension signal-to-noise ratio is detected to be no less than the minimum value of the historical feature dimension signal-to-noise ratio range, the execution is stopped. After the feature dimension signal-to-noise ratio optimization mechanism is completed, if the feature dimension signal-to-noise ratio is still less than the minimum value of the historical feature dimension signal-to-noise ratio range, a feature removal prompt will be sent; otherwise, the feature dimension adjusted by the sliding window will be marked as a qualified feature dimension and synchronized to the feature database for storage.

2. The personalized AI voiceprint cloning method as described in claim 1, characterized in that, The validity determination includes the following steps: Based on the signal conversion frequency obtained during the speech signal conversion process, it is compared with a set matching interval to locate the processing method corresponding to the interval to which it belongs. Specifically: If the acquired signal conversion frequency is within the first matching interval, the speech signal conversion process is deemed qualified, and the voiceprint features are extracted using the Mel frequency cepstral coefficient algorithm. If the acquired signal conversion frequency is within the second matching range, it is determined that the voice signal conversion process is distorted, and a re-acquisition warning is triggered to prompt the preset personnel to check the corresponding voice environment; If the acquired signal conversion frequency is within the third matching interval, it is determined that the speech signal conversion process is redundant, and the high-frequency components in the third matching interval are filtered out by a digital low-pass filter. The first matching interval, the second matching interval, and the third matching interval sequentially constitute the complete coverage range of the signal conversion frequency; Based on the positioning results of the matching interval, the preprocessing optimization of the converted speech signal is realized, and a signal quality confirmation instruction is sent. The signal quality confirmation instruction is used to prompt the preset personnel to review the digital speech data after the speech signal conversion in combination with the characteristics of the speech scene. If the review confirms that the core features of the speech are accurately preserved, the first conversion and optimization process is completed and the usability of the speech data is predicted; otherwise, a speech data collection warning is issued.

3. The personalized AI voiceprint cloning method as described in claim 2, characterized in that, The extraction of voiceprint features includes the following steps: After the speech signal conversion process is qualified, the target speech signal in the corresponding digital speech data is obtained and segmented into overlapping short frames with fixed duration. The overlapping short frames are used to approximate the continuously dynamically changing acoustic features in the target speech signal as short-time stationary signals. A windowing operation is performed on the segmented overlapping short frames to obtain the short frame speech signal of the corresponding overlapping short frame after the windowing operation, and a Fourier transform is performed. The Fourier transform is used to convert the time domain signal in the short frame speech signal into the frequency domain signal to obtain the power spectrum of the corresponding short frame speech signal. The windowing operation is used to reduce the spectral leakage at both ends of the overlapping short frames. The frequency axis of the power spectrum is introduced into the Mel frequency scale for filtering, and the Mel frequency cepstral coefficients are extracted to complete the initial extraction of voiceprint features. After the initial extraction of voiceprint features, the voiceprint features are obtained and then filtered to obtain voiceprint feature data.

4. The personalized AI voiceprint cloning method as described in claim 1, characterized in that, The voiceprint feature conversion includes the following steps: Qualified feature dimensions are obtained to determine speech speed adjustment and speech pitch adjustment. The speech speed adjustment is used to simulate the natural expression rhythm of the target voiceprint in different contexts based on the syllable duration in the voiceprint feature data. The speech pitch adjustment is used to determine the recognizability of the voiceprint feature based on the fundamental frequency information in the voiceprint feature data after speech speed adjustment. The specific process for determining the voice speed adjustment includes: L1, based on syllable speech rate and average speech rate, makes the following judgments: Based on the matching process between syllable duration in the voiceprint feature data and the target text, the syllable speech rate is obtained. At the same time, combined with the divided voiceprint feature transition time period, the average speech rate is obtained. The syllable duration represents the duration of a single syllable corresponding to the target voiceprint feature from the start to the end of pronunciation within the voiceprint feature transition time period. The syllable speech rate represents the pronunciation rate of a single syllable corresponding to the target voiceprint feature within the voiceprint feature transition time period. The average speech rate represents the average pronunciation rate of all syllables of the target voiceprint feature within the voiceprint feature transition time period. If the obtained syllable speech rate is within the first multiple of the average speech rate, the syllable speech rate is deemed to be qualified and no speech rate adjustment is required. If the obtained syllable speech speed is less than twice the average speech speed, it is determined that the syllable speech speed is slow, and the time axis stretching ratio is increased to slow down the cloned speech speed. If the obtained syllable speech speed is higher than the third multiple of the average speech speed, it is determined that the syllable speech speed is accelerated, and the time axis stretching ratio is reduced to accelerate the cloned speech speed. The first multiple, the second multiple, and the third multiple constitute a stepped threshold system for speech speed determination, and the values ​​increase sequentially. L2, Verification of voice speed adjustment effect: After adjusting the speech speed, the syllable speech speed and average speech speed are reacquired, and the corresponding ratio is obtained. If the ratio is within the first multiple range, the speech speed adjustment is deemed qualified and the current time axis stretching ratio is retained. At the same time, the speech pitch adjustment is determined. If the ratio is lower than the second multiple or higher than the third multiple, repeat the corresponding stretching ratio adjustment operation until the ratio is within the first multiple range; L3, Voice Tone Adjustment Judgment: The fundamental frequency difference is obtained based on the voiceprint feature data after speech speed adjustment. If the fundamental frequency difference is less than the set fundamental frequency difference, the speech tone is adjusted. Otherwise, the target voiceprint feature is obtained based on the result of voiceprint feature conversion, and a personalized AI voiceprint clone is completed to improve the recognition of voiceprint features. The fundamental frequency difference represents the difference between the maximum and minimum values ​​of the fundamental frequency of the voiceprint feature data.

5. The personalized AI voiceprint cloning method as described in claim 4, characterized in that, The voice tone adjustment includes the following steps: Frequency domain analysis is performed based on the acquired raw voiceprint feature data to obtain the energy value at each frequency point; The total energy of each frequency band is obtained based on the acquired frequency band range, which includes the first frequency band, the second frequency band, and the third frequency band; The total energy of each frequency band is compared with the set frequency band energy template to determine the energy ratio difference of each frequency band and send a frequency band energy ratio calibration command. The set frequency band energy template includes a set first frequency band, a second frequency band, and a third frequency band.

6. The personalized AI voiceprint cloning method as described in claim 5, characterized in that, The energy ratio difference of each frequency band includes the energy ratio difference of the first frequency band, the energy ratio difference of the second frequency band, and the energy ratio difference of the third frequency band; The energy ratio difference of the first frequency band represents the difference between the energy ratio of the first frequency band in the frequency range and the preset energy ratio in the first frequency band. The energy ratio difference of the second frequency band represents the difference between the energy ratio of the second frequency band in the frequency range and the preset energy ratio in the second frequency band. The energy ratio difference of the third frequency band represents the difference between the energy ratio of the third frequency band in the frequency band range and the preset energy ratio in the third frequency band. The frequency band energy ratio calibration command is used to adjust the energy ratio of each frequency band in a targeted manner based on the energy ratio difference between the first frequency band, the second frequency band, and the third frequency band.

7. A system applying the personalized AI voiceprint cloning method as described in any one of claims 1-6, characterized in that, include: The module includes a language signal conversion validity determination module, a voiceprint feature data acquisition and processing module, and a personalized voiceprint feature matching module. The language signal conversion validity determination module is used to collect the voice signal of the identified user through the microphone, convert the collected voice signal into digital voice data, and at the same time determine the validity of the voice signal conversion process based on the signal conversion frequency to ensure the accuracy of subsequent voiceprint feature extraction and cloning. The voiceprint feature data acquisition and processing module is used to collect the converted digital speech data, extract voiceprint features through the Mel frequency cepstral coefficient algorithm to obtain voiceprint feature data, and perform voiceprint feature processing. The personalized voiceprint feature matching module is used to convert voiceprint features based on the results of voiceprint feature processing, and combine the converted voiceprint feature data with the set speech text to complete the personalized AI voiceprint cloning.

8. A computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by a processor using the personalized AI voiceprint cloning method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • A Chinese text to personalized speech conversion method and system

    CN115240630B

  • A sound cloning method and system based on artificial intelligence

    CN117672182B

  • Personalized speech synthesis method, device and system and storage medium

    CN117219048A

  • Voice generation method and device, equipment and medium

    CN120279883A