Method and device for selecting reference audio for tone cloning

By integrating a comprehensive scoring mechanism that combines audio quality and text rhythm matching, a multi-dimensional audio quality evaluation system is constructed. This addresses the shortcomings of existing timbre cloning systems in selecting reference audio, enabling efficient and automated audio screening. It also improves the fidelity and naturalness of timbre transfer in the timbre cloning system, making it suitable for Chinese and English timbre cloning tasks.

CN120977285APending Publication Date: 2025-11-18BEIJING AISHU WISDOM TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511175348.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing voice cloning systems cannot accurately adapt to the synthesis requirements of different texts when selecting reference audio, resulting in synthesized speech that is difficult to achieve ideal results in terms of naturalness, fluency, and semantic consistency. In addition, they suffer from problems such as strong reliance on human intervention, limited selection dimensions, and inability to adapt to Chinese and English texts.

Method used

A comprehensive scoring mechanism that integrates audio quality and text rhythm matching is adopted. By evaluating the basic attributes of the audio and its matching degree with the target text, a multi-dimensional audio quality evaluation system is constructed, which is compatible with the differences in language structure between Chinese and English, and realizes automated, multi-dimensional selection of reference audio.

Benefits of technology

It improves the fidelity, stability and naturalness of timbre transfer, ensures that the selected reference audio is highly consistent with the target task in terms of structure, content and rhythm, reduces the problems of timbre drift and intonation abruptness, is suitable for Chinese and English dual-voice timbre cloning tasks, and has a lightweight and model-free system design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977285A_ABST
    Figure CN120977285A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a reference audio selection method and device for tone cloning, and the method comprises the following steps: obtaining an audio candidate set which comprises a plurality of candidate audios; extracting a basic attribute of each candidate audio, and determining an audio quality score of each candidate audio according to the basic attribute; according to a target text to be subjected to timbre cloning, evaluating a matching degree of each candidate audio and the target text in rhythm, and obtaining a text matching degree score of each candidate audio; and selecting one candidate audio from the audio candidate set as a reference audio according to the audio quality score and the text matching degree score of each candidate audio. According to the audio quality score and the text matching degree score of each candidate audio in the audio candidate set, the reference audio which is matched with the target text rhythm, good in quality and natural in structure is screened out, and therefore the fidelity, the stability and the naturalness of tone migration are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and specifically relates to a method and apparatus for selecting reference audio for timbre cloning. Background Technology

[0002] Voice cloning technology aims to reproduce the vocal characteristics of a target speaker, enabling any text content to be presented as the voice of a specified speaker. Such systems typically require reference audio of the speaker as input, serving as the basis for vocal feature extraction and transfer.

[0003] In typical voice cloning tasks, each speaker often has multiple speech samples, from which the system needs to select the most suitable audio sample as a reference voice. Because these audio samples vary significantly in content, duration, speech rate, noise level, and naturalness, inappropriate selection will significantly affect the stability and realism of the final voice cloning result.

[0004] Currently, some timbre cloning systems use simple rule matching when selecting audio, such as filtering based solely on audio duration or randomly selecting audio samples; some systems are based on speaker embedding methods, extracting speaker feature vectors from audio through neural networks and then matching target timbre samples based on vector similarity; and some are based on reference audio methods, directly using a reference audio segment provided by the user as timbre input.

[0005] However, simple rule matching cannot accurately adapt to the synthesis needs of different texts; speaker embedding-based methods require additional feature extraction models, which have a large computational cost; in reference audio-based methods, the quality and length of the reference audio will significantly affect the synthesis effect, making it difficult for the synthesized speech to achieve ideal results in terms of naturalness, fluency and semantic consistency.

[0006] Application content

[0007] The purpose of this application is to provide a method and apparatus for selecting reference audio for timbre cloning, so as to solve the defect of the prior art that it cannot select ideal reference audio.

[0008] To solve the above-mentioned technical problems, this application is implemented as follows:

[0009] Firstly, a method for selecting reference audio for timbre cloning is provided, comprising the following steps:

[0010] Obtain an audio candidate set, which includes multiple candidate audio files;

[0011] Extract the basic attributes of each candidate audio, and determine the audio quality score of each candidate audio based on the basic attributes;

[0012] Based on the target text to be cloned, the rhythmic matching degree between each candidate audio and the target text is evaluated to obtain a text matching score for each candidate audio.

[0013] Based on the audio quality score and text matching score of each candidate audio, a candidate audio is selected from the audio candidate set as a reference audio.

[0014] Secondly, a selection device for reference audio for timbre cloning is provided, comprising:

[0015] The acquisition module is used to acquire an audio candidate set, which includes multiple candidate audio files;

[0016] An extraction module is used to extract the basic attributes of each candidate audio and determine the audio quality score of each candidate audio based on the basic attributes.

[0017] The evaluation module is used to evaluate the rhythmic matching degree between each candidate audio and the target text to be cloned, and to obtain a text matching score for each candidate audio.

[0018] The selection module is used to select a candidate audio as a reference audio from the audio candidate set based on the audio quality score and text matching score of each candidate audio.

[0019] In this embodiment of the application, reference audio that matches the rhythm of the target text, is of good quality, and has a natural structure is selected based on the audio quality score and text matching score of each candidate audio in the audio candidate set, thereby improving the fidelity, stability and naturalness of timbre transfer. Attached Figure Description

[0020] Figure 1 This is a flowchart of a method for selecting reference audio for timbre cloning provided in an embodiment of this application;

[0021] Figure 2 This is a specific implementation diagram of the reference audio selection method provided in the embodiments of this application;

[0022] Figure 3 This is a schematic diagram of a reference audio selection device for timbre cloning provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] In timbre cloning technology, an intelligent, stable, and automatically executed audio selection mechanism is needed to select the audio with the best quality and most matching match to the target text from multiple candidate speech samples as a timbre reference without human intervention, so as to improve the fidelity of timbre transfer and the naturalness of expression.

[0025] However, existing technologies do not adequately consider factors such as the matching degree between text and audio lengths and the relevance between audio content and the text to be synthesized; they lack comprehensive analysis of audio quality and content, which may result in the selection of samples with noise, unclear pronunciation, or abnormal intonation; and they do not differentiate between the differences in the structure of Chinese and English languages, which can easily lead to mismatches in cross-language tasks.

[0026] To address the above shortcomings, this application provides a method for selecting reference audio for a Chinese-English voice cloning system. This method automatically selects reference samples from multiple candidate speaker audios that match the rhythm of the target text, are of high quality, and have a natural structure, thereby improving the fidelity, stability, and naturalness of voice transfer. This enhances the quality of speech synthesis, resulting in better performance in terms of naturalness, fluency, and semantic consistency, thus resolving the problem of inconsistent speech synthesis results in existing methods.

[0027] Specifically, this application addresses the problems of existing timbre cloning systems in the reference audio selection process, such as strong reliance on manual intervention, limited selection dimensions, and inability to adapt to Chinese and English text. It proposes an automated, multi-dimensional, and language-structure-adaptive method for selecting optimal reference audio, with the following technical innovations:

[0028] 1. A comprehensive scoring mechanism integrating audio quality and text rhythm matching is proposed. The system not only evaluates the basic attributes and content quality of the audio, but also estimates the reading duration of the target text and compares it with the audio duration. By weighted fusion of quality score and matching score, audio selection that better meets the needs of expression rhythm is achieved.

[0029] 2. A multi-dimensional audio quality evaluation system was constructed, covering five key aspects: reasonable audio duration, background noise level, natural expression, emotional stability, and pronunciation clarity. Compared with traditional methods that only consider duration or rely on speech embedding similarity, the embodiments of this application more comprehensively measure the usability and transfer value of audio samples.

[0030] 3. To accommodate the differences in language structure between Chinese and English, a language unit weighting mechanism is introduced in the text rhythm estimation stage, effectively solving the adaptation problem caused by the different expression densities of Chinese characters and English words.

[0031] 4. It possesses model independence and system scalability, and can independently complete high-quality audio screening tasks without relying on neural network embedding models, adapting to the deployment requirements of multi-speaker, multi-language, and multi-task timbre cloning systems.

[0032] The following description, in conjunction with the accompanying drawings, details a method for selecting reference audio for timbre cloning provided in this application through specific embodiments and application scenarios.

[0033] like Figure 1 The diagram shown is a flowchart of a method for selecting reference audio for timbre cloning according to an embodiment of this application. The method includes the following steps:

[0034] Step 101: Obtain an audio candidate set, which includes multiple candidate audio files.

[0035] Step 102: Extract the basic attributes of each candidate audio and determine the audio quality score of each candidate audio based on the basic attributes.

[0036] Specifically, the duration rationality, background noise level, speech naturalness, emotional stability, and pronunciation clarity of each candidate audio can be extracted.

[0037] Step 103: Based on the target text to be cloned, evaluate the rhythmic matching degree between each candidate audio and the target text to obtain a text matching score for each candidate audio.

[0038] Specifically, the theoretical reading duration of the target text can be calculated based on its language type and language structure features. The audio duration of each candidate audio is then compared with the theoretical reading duration. Based on the difference in duration, the rhythmic matching degree between the candidate audio and the target text is evaluated, and a text matching score for each candidate audio is obtained.

[0039] Step 104: Select a candidate audio from the audio candidate set as a reference audio based on the audio quality score and text matching score of each candidate audio.

[0040] Specifically, a weighted fusion can be performed based on the audio quality score and text matching score of each candidate audio to calculate the final comprehensive score of each candidate audio; the candidate audio with the highest final comprehensive score is selected from the audio candidate set as a reference audio for use by the timbre cloning model.

[0041] In this embodiment, the process and results of selecting reference audio can also be recorded in log and mapping files to ensure that the entire reasoning process is traceable and reconfigurable.

[0042] In this embodiment of the application, reference audio that matches the rhythm of the target text, is of good quality, and has a natural structure is selected based on the audio quality score and text matching score of each candidate audio in the audio candidate set, thereby improving the fidelity, stability and naturalness of timbre transfer.

[0043] This application proposes an automatic reference audio selection method for a Chinese-English voice cloning system. The aim is to select the most suitable high-quality audio samples for reference from a speaker audio resource library without manual intervention, thereby ensuring the voice cloning system's voice fidelity and naturalness during speech generation. Figure 2 As shown, this method mainly includes seven core steps: candidate audio information extraction, target text acquisition, target text rhythm estimation, audio quality assessment, text-audio matching degree assessment, comprehensive scoring and ranking, and result output and recording. The following provides a detailed explanation of each step:

[0044] 1. Audio candidate set acquisition

[0045] First, the system scans a directory of specified speaker voice data to extract all available audio files as candidate timbre resources. The system allows setting filtering criteria to retain only audio files that meet the format requirements and records their basic metadata, including file path, speaker identifier, file name, and audio duration. The goal of this step is to establish a complete and analyzable audio candidate set, providing a data foundation for subsequent quality assessment and matching scoring. In practical deployments, this step supports batch processing and is adaptable to scenarios with multiple speakers and large-scale timbre databases.

[0046] 2. Obtaining the text to be synthesized

[0047] The system reads the file containing the target text from a specified path. Each line of text is treated as an independent synthesis task, loaded sequentially, and a processing list is created for use by the subsequent audio matching module.

[0048] 3. Text length estimation

[0049] For the target text to be cloned, the system first identifies its language type (e.g., Chinese or English) and calculates its theoretical reading time based on language structure features. This calculation is based on a weighted length model of language units: for Chinese text, each Chinese character is counted as 1.5 units of length; for English text, each word is counted as 1.0 unit of length. The system estimates the reading time of the entire text based on a preset speech rate estimate (4 units of length per second), serving as a reference value for rhythm matching. This step is crucial for ensuring that the selected audio matches the text expression in terms of speech rate and rhythm.

[0050] 4. Audio quality rating

[0051] The audio quality assessment module proposed in this application mainly performs detailed quantitative scoring on candidate audio samples from five dimensions: duration reasonableness, background noise level, speech naturalness, emotional stability, and pronunciation clarity. Each dimension corresponds to a specific detection mechanism and scoring standard, ensuring that the system can comprehensively screen out high-quality reference audio suitable for timbre cloning.

[0052] Reasonableness of duration:

[0053] The duration of candidate reference audio is one of the important factors affecting the timbre cloning effect. If the duration is too short (e.g., less than 2 seconds), it often cannot provide sufficient timbre feature information, which can easily lead to insufficient model learning and incomplete timbre transfer; while if the duration is too long (e.g., more than 20 seconds), it may contain multiple semantic segments, topic changes or intonation shifts, which increases the complexity of timbre model extraction and affects consistency and generalization ability.

[0054] This application sets a reasonable reference audio duration range of 3 to 15 seconds. Audio within this range receives a full score (e.g., 100 points). If the audio duration is shorter than 3 seconds, the system deducts 1 point for every 0.1 seconds shorter; if it is longer than 15 seconds, 1 point is deducted for every 0.5 seconds longer. For obviously abnormal audio (e.g., less than 1 second or more than 30 seconds), the system will directly remove it from the candidate set and it will not participate in subsequent scoring. This scoring mechanism balances the integrity of timbre expression with computational efficiency, avoiding inefficient samples from interfering with model judgment.

[0055] Background noise level:

[0056] The system analyzes the purity of audio signals using methods such as signal-to-noise ratio (SNR) estimation, silent zone detection, and spectrum analysis. A higher SNR indicates clearer vocals in the audio. The system sets an SNR threshold of 20dB; signals below this threshold are penalized accordingly. Simultaneously, the system also detects noise peaks, low-frequency AC interference, and background noise (such as wind, vocals, and keyboard sounds) in silent sections of the audio, adjusting the score accordingly. The scoring results are divided into multiple levels; an SNR above 30dB receives full marks, while an SNR below 10dB is considered unacceptable and is rejected.

[0057] Naturalness rating:

[0058] The system uses endpoint detection, voice activity detection (VAD), and inter-frame energy fluctuation analysis to identify whether there are artificial segmentation marks, rhythmic breaks, or unnatural interruptions in the speech. The system also detects rhythmic smoothness, such as whether there are sudden stops, pauses, or segment jumps. A high naturalness score is awarded if the audio is continuous, fluent, and without abrupt stops or jumps; points are deducted for disjointed speech segments or sudden energy fluctuations.

[0059] Emotional stability:

[0060] The system uses fundamental frequency (F0) trajectory analysis and an emotion feature recognition model to quantitatively evaluate emotional expression in speech. If the audio contains obvious strong emotional expressions such as anger, surprise, sobbing, laughter, or a recitation-like tone, the system classifies it as a non-neutral expression and lowers its score. To ensure the universality and adaptability of the cloned speech, this method prioritizes selecting reference samples with stable intonation and neutral emotion. If the entire audio is in a highly motivated or exaggerated state, it may be rejected by the system.

[0061] Pronunciation clarity:

[0062] The system analyzes whether syllables are complete, whether there are any unclear sounds, swallowed sounds, or glissando caused by excessively fast pronunciation, and whether pauses are even and reasonable. The system uses techniques such as formant trajectory recognition, speech-noise separation, and energy envelope analysis to judge pronunciation quality, ensuring that selected reference samples have clear speech and well-defined structure. Abnormal pronunciations such as stuttering, coughing, or excessively loud breathing sounds are also penalized and included in the deduction mechanism.

[0063] 5. Audio length matching score

[0064] After estimating the text rhythm and extracting the audio duration, the system compares and analyzes the two to evaluate the rhythmic matching degree between the audio sample and the target text. Specifically, it calculates the duration difference between the two and assigns different matching scores based on preset scoring ranges. For example, if the difference between the audio duration and the estimated text duration is within 1 second, the score is 100 points; between 1 and 3 seconds, 80 points; between 3 and 5 seconds, 60 points; between 5 and 10 seconds, 30 points; and more than 10 seconds, only 10 points. This step effectively eliminates audio with excessively large differences in speech rate or mismatched rhythmic styles, improving the consistency of subsequent timbre transfer.

[0065] 6. Comprehensive score calculation and ranking

[0066] After scoring the audio quality and text matching accuracy, the system weights and fuses these two scores to calculate a final composite score for each candidate audio file. The weighting formula is as follows:

[0067] S final =α·S 音频质量 +(1-α)·S 文本匹配度

[0068] Here, α represents the weight of the audio quality score, with a default setting of 0.6, emphasizing that the quality of the audio itself is more important than rhythm matching. The system sorts all candidate audios in descending order based on their final scores, prioritizing the audio with the highest score as the reference sample. If multiple audios have the same score, the system may prioritize the shorter sample to reduce the subsequent processing burden and improve system response efficiency while ensuring quality.

[0069] 7. Optimal Audio Selection and Output

[0070] Finally, the system sorts candidates based on their overall scores and automatically selects the highest-scoring audio file from the candidate set as the reference timbre input for the text, which is then used by the timbre cloning model. Simultaneously, this selection process and results are recorded in logs and mapping files to ensure the traceability and reconstructability of the entire inference process. If, during the screening process, all candidate audio files are found to be unsatisfactory, the system will automatically output a prompt message, guiding the user to provide or supplement the audio resources.

[0071] Through the above steps, the embodiments of this application can automatically and stably complete the selection of audio samples without human intervention, providing speaker prompt audio with structural matching and reliable quality for speech synthesis, thereby improving the overall performance and practical usability of the speech synthesis system.

[0072] The automatic reference audio optimization method provided in this application embodiment shows significant effects in the timbre cloning system, mainly in the following three aspects:

[0073] First, in terms of timbre transfer quality, by introducing a dual scoring mechanism of audio quality and text matching, we ensure that the selected reference audio is highly consistent with the target task in multiple dimensions such as structure, content, and rhythm, effectively reducing problems such as timbre drift and abrupt changes in intonation caused by inconsistent quality of reference samples.

[0074] Second, in terms of adaptability, it innovatively adopts a weighted modeling mechanism for Chinese characters and English words to achieve unified rhythm estimation and matching of Chinese and English texts, so that the selection method has good generalization ability in multilingual environments and is suitable for Chinese and English dual-voice chromaticity cloning tasks.

[0075] Third, in terms of resource efficiency, the system design follows the principles of lightweight and model-free operation. All scoring mechanisms are based on rule-based and statistical calculations, resulting in fast and stable computation, making it suitable for large-scale deployment. This method, as an independent module, is easy to integrate and can be applied to the preprocessing workflow of existing timbre cloning systems, effectively improving overall processing efficiency and result stability.

[0076] like Figure 3 The diagram shown is a schematic representation of a reference audio selection device for timbre cloning provided in an embodiment of this application, comprising:

[0077] The acquisition module 310 is used to acquire an audio candidate set, which includes multiple candidate audios.

[0078] The extraction module 320 is used to extract the basic attributes of each candidate audio and determine the audio quality score of each candidate audio based on the basic attributes.

[0079] Specifically, the extraction module 320 is used to extract the duration reasonableness, background noise level, speech naturalness, emotional stability and pronunciation clarity of each candidate audio.

[0080] Evaluation module 330 is used to evaluate the degree of rhythmic matching between each candidate audio and the target text based on the target text to be cloned, and to obtain a text matching score for each candidate audio.

[0081] Specifically, the evaluation module 330 is used to calculate the theoretical reading duration of the target text based on the language type and language structure features of the target text to be cloned; compare and analyze the audio duration of each candidate audio with the theoretical reading duration, evaluate the degree of rhythm matching between the candidate audio and the target text based on the duration difference, and obtain a text matching score for each candidate audio.

[0082] Selection module 340 is used to select a candidate audio as a reference audio from the audio candidate set based on the audio quality score and text matching score of each candidate audio.

[0083] Specifically, the selection module 340 is used to perform weighted fusion based on the audio quality score and text matching score of each candidate audio to calculate the final comprehensive score of each candidate audio; and select the candidate audio with the highest final comprehensive score from the audio candidate set as the reference audio for use by the timbre cloning model.

[0084] In this embodiment, the above-mentioned device further includes:

[0085] The recording module is used to record the process and results of selecting reference audio to log and mapping files, ensuring that the entire inference process is traceable and reconfigurable.

[0086] In this embodiment of the application, reference audio that matches the rhythm of the target text, is of good quality, and has a natural structure is selected based on the audio quality score and text matching score of each candidate audio in the audio candidate set, thereby improving the fidelity, stability and naturalness of timbre transfer.

[0087] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described method embodiment for selecting reference audio for timbre cloning, and achieves the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0088] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0089] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0090] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for selecting reference audio for timbre cloning, characterized in that, Includes the following steps: Obtain an audio candidate set, which includes multiple candidate audio files; Extract the basic attributes of each candidate audio, and determine the audio quality score of each candidate audio based on the basic attributes; Based on the target text to be cloned, the rhythmic matching degree between each candidate audio and the target text is evaluated to obtain a text matching score for each candidate audio. Based on the audio quality score and text matching score of each candidate audio, a candidate audio is selected from the audio candidate set as a reference audio.

2. The method according to claim 1, characterized in that, The extraction of the basic attributes of each candidate audio specifically includes: Extract the duration reasonableness, background noise level, speech naturalness, emotional stability and pronunciation clarity of each candidate audio.

3. The method according to claim 1, characterized in that, The step of evaluating the rhythmic matching degree between each candidate audio and the target text to be cloned, based on the target text, to obtain a text matching score for each candidate audio, specifically includes: Based on the language type and language structure characteristics of the target text to be cloned, calculate the theoretical reading time of the target text; The audio duration of each candidate audio is compared and analyzed with the theoretical reading duration. Based on the difference in duration, the degree of rhythmic matching between the candidate audio and the target text is evaluated, and a text matching score for each candidate audio is obtained.

4. The method according to claim 1, characterized in that, The step of selecting a candidate audio as a reference audio from the audio candidate set based on the audio quality score and text matching score of each candidate audio specifically includes: The final comprehensive score for each candidate audio is calculated by weighting and fusing the audio quality score and text matching score of each candidate audio. The candidate audio with the highest final comprehensive score is selected from the audio candidate set as the reference audio for use by the timbre cloning model.

5. The method according to claim 1, characterized in that, Also includes: The process and results of selecting reference audio are recorded in log and mapping files to ensure that the entire inference process is traceable and reconfigurable.

6. A device for selecting reference audio for timbre cloning, characterized in that, include: The acquisition module is used to acquire an audio candidate set, which includes multiple candidate audio files; An extraction module is used to extract the basic attributes of each candidate audio and determine the audio quality score of each candidate audio based on the basic attributes. The evaluation module is used to evaluate the rhythmic matching degree between each candidate audio and the target text to be cloned, and to obtain a text matching score for each candidate audio. The selection module is used to select a candidate audio as a reference audio from the audio candidate set based on the audio quality score and text matching score of each candidate audio.

7. The apparatus according to claim 6, characterized in that, The extraction module is specifically used to extract the duration rationality, background noise level, speech naturalness, emotional stability, and pronunciation clarity of each candidate audio.

8. The apparatus according to claim 6, characterized in that, The evaluation module is specifically used to calculate the theoretical reading duration of the target text based on its language type and language structure features; compare the audio duration of each candidate audio with the theoretical reading duration; evaluate the rhythmic matching degree between the candidate audio and the target text based on the duration difference between the two; and obtain a text matching score for each candidate audio.

9. The apparatus according to claim 6, characterized in that, The selection module is specifically used to perform weighted fusion based on the audio quality score and text matching score of each candidate audio to calculate the final comprehensive score of each candidate audio; and to select the candidate audio with the highest final comprehensive score from the audio candidate set as the reference audio for use by the timbre cloning model.

10. The apparatus according to claim 6, characterized in that, Also includes: The recording module is used to record the process and results of selecting reference audio to log and mapping files, ensuring that the entire inference process is traceable and reconfigurable.

Citation Information

Patent Citations

  • Synthesis by generation and concatenation of multi-form segments

    CN101828218A

  • Method and device for voice synthesis for Chinese teaching

    CN102723077A

  • Timbre selection method and device, electronic equipment, readable storage medium and program product

    CN116110366A

  • Audio resume generation and page access method and device, storage medium and program product

    CN119132334A

  • Mediator timbre cloning method and system, electronic equipment and storage medium

    CN120089124A