Accent conversion method and device based on discrete speech representation, equipment and medium

By using discrete speech representation, the speech to be converted is discretized and timbre features are synthesized, which solves the problem of insufficient accuracy and naturalness in accent conversion in the existing technology and achieves high accuracy and naturalness in accent conversion.

CN121811896APending Publication Date: 2026-04-07PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing accent conversion technologies struggle to collect reference speech for the target accent in practical applications, leading to decreased speech accuracy and naturalness, as well as numerous synthetic artifacts.

Method used

A discrete speech representation-based approach is adopted to discretize the speech to be converted, extract the initial speech unit sequence and original accent features, perform autoregressive conversion using a preset accent conversion model, and synthesize the target speech by combining timbre features.

Benefits of technology

It improves the accuracy and naturalness of accent conversion, removes redundant features, preserves the speaker's timbre, and enhances the natural fluency of speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811896A_ABST
    Figure CN121811896A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, and particularly discloses an accent conversion method and device based on discrete voice representation, equipment and a medium. According to the method, the to-be-converted voice is discretized to obtain the voice unit which only retains phoneme information, and then the voice unit is mapped into the target accent by using the accent conversion model, so that redundant features are removed, the model is enabled to concentrate on accent conversion, and the accuracy of accent conversion is improved; and secondly, the target voice is synthesized in combination with the tone characteristics of the speaker, so that the tone of the speaker is reserved, and the naturalness of the target voice after accent conversion is improved. The method is applied to financial services such as telemarketing, telephone banking and telephone return visit and medical systems such as remote inquiry, patient follow-up visit and medical education, the voice accent of a speaker in communication participants can be accurately converted into a target accent familiar with a listener, the tone of the speaker is kept, and the communication efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to an accent conversion method, apparatus, device and medium based on discrete speech representation. Background Technology

[0002] With the deepening of digital transformation, voice interaction technology has become a key infrastructure for improving service efficiency and optimizing user experience in the financial and healthcare industries. Voice interaction is widely used in financial services such as telephone sales, telephone banking, and telephone follow-ups, as well as in healthcare systems such as remote consultations, patient follow-ups, and medical education. Financial and medical services cover all population groups. For customers with strong dialects or speech impairments, voice accuracy is a core requirement for voice interaction. For example, in doctor-patient communication, mishearing or misunderstanding of key information such as symptoms, medical history, and medication dosage can directly lead to misdiagnosis. Therefore, converting the speaker's accent into the target accent to avoid mishearing and misunderstanding is a crucial step in voice interaction.

[0003] Existing accent conversion or accent normalization techniques require target accent reference speech for practical inference applications. However, target accent reference speech is often difficult to collect and its quality varies greatly. Furthermore, synthesis artifacts are frequently introduced during the conversion process, leading to a decrease in speech accuracy and naturalness. Therefore, improving the accuracy and naturalness of accent conversion has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides a method, apparatus, device, and medium for accent conversion based on discrete speech representation, in order to improve the accuracy and naturalness of accent conversion.

[0005] In a first aspect, this application provides an accent conversion method based on discrete speech representation, the method comprising: Upon receiving the speech to be converted, the speech is discretized and accent extracted to obtain a discretized initial speech unit sequence and original accent features. Based on a preset accent conversion model, an autoregressive conversion is performed on the initial speech unit sequence and the original accent features to obtain the target speech unit sequence corresponding to the target accent. The timbre features corresponding to the speech to be converted are obtained, and the target speech corresponding to the target accent is synthesized based on the timbre features and the target speech unit sequence.

[0006] Secondly, this application also provides an accent conversion device based on discrete speech representation, the device comprising: The speech discretization module is used to discretize and extract accents from the received speech to be converted, thereby obtaining a discretized initial speech unit sequence and original accent features. The target accent conversion module is used to perform autoregressive conversion on the initial speech unit sequence and the original accent features based on a preset accent conversion model to obtain the target speech unit sequence corresponding to the target accent. The target speech acquisition module is used to acquire the timbre features corresponding to the speech to be converted, and synthesize the target speech corresponding to the target accent based on the timbre features and the target speech unit sequence.

[0007] Thirdly, this application also provides a computer device, the computer device including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the accent conversion method based on discrete speech representation as described above.

[0008] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the above-described accent conversion method based on discrete speech representation.

[0009] This application discloses an accent conversion method, apparatus, device, and medium based on discrete speech representation. Upon receiving speech to be converted, the method discretizes and extracts the accent, obtaining a discretized initial speech unit sequence and original accent features. Based on a preset accent conversion model, an autoregressive transformation is performed on the initial speech unit sequence and the original accent features to obtain a target speech unit sequence corresponding to the target accent. The method also obtains the timbre features corresponding to the speech to be converted, and synthesizes the target speech corresponding to the target accent based on the timbre features and the target speech unit sequence. This application discretizes the speech to be converted to obtain speech units that retain only phoneme information. Then, it uses an accent conversion model to map the speech units to the target accent, removing redundant features and allowing the model to focus on accent conversion, thus improving the accuracy of accent conversion. Furthermore, by combining the speaker's timbre features with the synthesized target speech, the method preserves the speaker's timbre, improving the naturalness of the target speech after accent conversion. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a first schematic flowchart of an accent conversion method based on discrete speech representation provided by an embodiment of this application; Figure 2 This is a flowchart of an accent conversion provided by an embodiment of this application; Figure 3 This is a second schematic flowchart of an accent conversion method based on discrete speech representation provided in an embodiment of this application; Figure 4 This is a third schematic flowchart of an accent conversion method based on discrete speech representation provided in an embodiment of this application; Figure 5 A schematic block diagram of an accent conversion device based on discrete speech representation provided for embodiments of this application; Figure 6 A schematic block diagram of the structure of a computer device provided for an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0014] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0015] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0016] This application provides an embodiment of an accent conversion method, apparatus, device, and medium based on discrete speech representation. The accent conversion method based on discrete speech representation can be applied to a server. It discretizes the speech to be converted to obtain speech units that retain only phoneme information. Then, an accent conversion model is used to map the speech units to the target accent, removing redundant features and allowing the model to focus on accent conversion, thus improving the accuracy of the accent conversion. Secondly, the target speech is synthesized by combining the speaker's timbre features, preserving the speaker's timbre and improving the naturalness of the target speech after accent conversion. The server can be a standalone server or a server cluster.

[0017] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0018] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating an accent conversion method based on discrete speech representation provided in an embodiment of this application. This accent conversion method based on discrete speech representation can be applied in a server to discretize the speech to be converted, obtaining speech units that retain only phoneme information. Then, an accent conversion model is used to map the speech units to the target accent, removing redundant features and allowing the model to focus on accent conversion, thus improving the accuracy of accent conversion. Secondly, the target speech is synthesized by combining the speaker's timbre features, preserving the speaker's timbre and improving the naturalness of the target speech after accent conversion.

[0019] like Figure 1 As shown, the accent conversion method based on discrete speech representation specifically includes steps S101 to S103.

[0020] S101. Upon receiving the speech to be converted, the speech to be converted is discretized and the accent is extracted to obtain a discretized initial speech unit sequence and original accent features. In one embodiment, the speech to be converted is speech with an accent category other than the target accent. For example, when an insurance customer service representative makes an insurance recommendation call, in order to convert the representative's accent into an accent that the customer can understand, it is necessary to perform accent conversion on the representative's speech.

[0021] In one embodiment, upon receiving the speech to be converted, preprocessing is performed on the speech. The preprocessing steps include: removing environmental noise (such as background noise and device interference) using spectral subtraction or a deep learning model to ensure clear speech signal; unifying the sampling rate and quantization bit depth, and normalizing the audio volume to eliminate the influence of hardware differences; extracting valid speech segments and excluding silent or non-speech parts (such as coughing and breathing sounds).

[0022] Further, the step of discretizing and extracting accents from the received speech to be converted to obtain a discretized initial speech unit sequence and original accent features includes: upon receiving the speech to be converted, discretizing the speech based on a preset self-supervised learning model to obtain discrete speech units, and clustering the discrete speech units to obtain the initial speech unit sequence; and extracting accents from the speech based on a preset accent editor to obtain the original accent features.

[0023] In one embodiment, the self-supervised learning model can be a pre-trained HuBERT model, such as... Figure 2 As shown, the HuberT model is used to encode the preprocessed speech to be converted, and the hidden states of the intermediate layers (such as the output of layer 6, dimension 768) are extracted. Discrete speech tokens are extracted using self-supervised models such as HuberT, mainly preserving phoneme-level information while weakening accent and prosodic features. Subsequently, an accent conversion model is used to perform accent autoregressive conversion on the tokens to achieve the mapping from non-target accents to the target accent, which can effectively reduce the dependence on the quality of synthesized speech. K-means clustering is performed on the continuous hidden states output by HuberT, mapping each speech frame to the nearest cluster center, resulting in a discrete speech token (unit) sequence, i.e., the initial speech unit sequence.

[0024] In one embodiment, such as Figure 2 As shown, the speech to be converted is transmitted to the accent encoder, which extracts the accent and obtains the original accent features.

[0025] In the above embodiments, phoneme-level speech features are obtained through self-supervised learning, and accent extraction is performed on the speech to be processed through an accent encoder, so that phoneme information and accent information are separated. Discretization processing ensures that phoneme content is not lost, and accent extraction accurately locates the accent parameters to be corrected, laying the foundation for converting non-target accents into standard accents in the future.

[0026] Furthermore, before discretizing and extracting the accent of the speech to be converted upon receiving the speech to be converted, and obtaining the discretized initial speech unit sequence and the original accent features, the method further includes: upon receiving the speech audio, performing accent recognition on the speech audio based on an accent classification model to determine the accent category corresponding to the speech audio; and identifying audio with an accent category other than the target accent as the speech to be converted.

[0027] In one embodiment, upon receiving voice audio, the voice audio is preprocessed to eliminate interference such as environmental noise and channel distortion, ensuring clear accent features.

[0028] Preprocessing may include: removing background noise (such as street noise and equipment current noise) using spectral subtraction or deep learning models; extracting effective speech segments and excluding silent or non-speech parts; and unifying the sampling rate and volume of the audio to avoid hardware differences affecting feature extraction.

[0029] In one embodiment, the accent classification model is trained using a deep learning model. The input parameter is the speech audio, and the output parameter is at least one accent category of the speech audio and the probability corresponding to each accent category.

[0030] Key acoustic features, including Mel frequency cepstral coefficients and spectral envelope, are extracted from speech audio using an accent classification model to effectively distinguish accents.

[0031] By comparing each key acoustic feature with the features corresponding to each accent, the similarity between each key acoustic feature and each feature corresponding to the accent is determined, and then the probability of the accent of the speech audio being of each category is obtained by weighted calculation.

[0032] The accent category with a probability exceeding a preset probability threshold is determined as the accent category of the currently input speech audio.

[0033] In one embodiment, if the accent category is not the target accent, the audio audio is identified as the speech to be converted.

[0034] In another embodiment, the audio can be individual voice audio or long conversation audio from a group, such as movie dubbing, medical consultation audio, or financial conference audio. When the audio is a long conversation, the accent classification model can identify the audio corresponding to the target accent and the audio corresponding to the non-target accent in the long conversation, and select the audio segment corresponding to the non-target accent as the audio to be converted. For example, medical consultations and financial conferences often have multiple participants from different regions, and the accents of each participant may differ or be the same. In order for each participant or the viewer of the meeting record to hear what the speaker is saying clearly, it is necessary to convert the speaker's accent into the target accent corresponding to each participant or a specified target accent. If the speaker's accent is the same as the target accent, no conversion is required; if the speaker's accent is different from the target accent, then the audio segment of that participant's voice is used as the accent to be converted.

[0035] In the above embodiments, the accent classification model accurately identifies accent categories and filters out non-target accents. In medical discussion meetings with multiple participants, it can accurately identify the accents of each participant's speech, determine the speech to be converted, and improve the efficiency of meeting communication.

[0036] S102. Based on a preset accent conversion model, perform autoregressive conversion on the initial speech unit sequence and the original accent features to obtain the target speech unit sequence corresponding to the target accent. In one embodiment, since tokens primarily capture phoneme category information, and accent differences are only reflected in the distribution differences of token sequences, accent normalization can be effectively achieved through token mapping. For example... Figure 2 As shown, an autoregressive transformation is performed on the initial speech unit sequence using a pre-trained accent conversion model, mapping the initial token sequence of the non-target accent to the token sequence of the target accent, i.e. the target speech unit sequence, thereby achieving phoneme-level accent normalization.

[0037] The accent conversion model can be structured as an encoder-decoder. The encoder extracts contextual features from the initial speech unit sequence and captures the temporal dependencies of the phoneme sequence. The decoder learns the mapping pattern between the initial speech unit sequence and the target accent tokens through an attention mechanism, and outputs the target accent token sequence, i.e., the target speech unit sequence.

[0038] In one embodiment, the accent conversion model employs a non-parallel training mechanism, where the training data consists of token sequences of non-standard speech (speech audio corresponding to the non-target accent) and token sequences of standard speech (speech audio corresponding to the target accent) synthesized by TTS (Text-to-Speech), without requiring content alignment.

[0039] S103. Obtain the timbre features corresponding to the speech to be converted, and synthesize the target speech corresponding to the target accent based on the timbre features and the target speech unit sequence.

[0040] In one embodiment, to preserve the original speaker's timbre, the target speech unit sequence is synthesized into natural and fluent target accent speech. The timbre features of the speech to be converted are obtained as a speaker embedding vector, specifically, as follows: Figure 2 As shown, a pre-trained speaker encoder is used to process the speech to be converted, outputting a fixed-dimensional speaker embedding vector that uniquely represents the speaker's timbre features. Speech synthesis is then performed by combining these timbre features with the target speech unit sequence, resulting in synthesized target speech that retains the speaker's timbre.

[0041] The above embodiments provide a method, apparatus, device, and medium for accent conversion based on discrete speech representation. Upon receiving speech to be converted, the speech is discretized and accent extracted to obtain a discretized initial speech unit sequence and original accent features. Based on a preset accent conversion model, an autoregressive transformation is performed on the initial speech unit sequence and the original accent features to obtain a target speech unit sequence corresponding to the target accent. The timbre features corresponding to the speech to be converted are obtained, and based on the timbre features and the target speech unit sequence, the target speech corresponding to the target accent is synthesized. This application discretizes the speech to be converted to obtain speech units that retain only phoneme information. Then, an accent conversion model is used to map the speech units to the target accent, removing redundant features and allowing the model to focus on accent conversion, thus improving the accuracy of accent conversion. Furthermore, the target speech is synthesized by combining the speaker's timbre features, preserving the speaker's timbre and improving the naturalness of the target speech after accent conversion.

[0042] Please see Figure 3 , Figure 3 This is a schematic flowchart illustrating an accent conversion method based on discrete speech representation provided in an embodiment of this application. This accent conversion method based on discrete speech representation can be applied to servers, uses non-parallel data to train the accent conversion model, can quickly generate large-scale standard speech, solves the problem of scarce parallel corpora, and reduces the cost of obtaining standard speech.

[0043] like Figure 3 As shown, the accent conversion method based on discrete speech representation includes steps S201 to S203 before step S102.

[0044] S201. Obtain at least one non-standard speech corresponding to the non-target accent and at least one standard speech corresponding to the target accent; In one embodiment, non-target accented speech is collected from public speech libraries or scenario-specific speech libraries (such as telephone customer service or language learner pronunciation samples), covering a variety of accent types (such as Chinglish, Indian English, Japanese English, etc.).

[0045] In a specific embodiment, for a speech library with multiple accent categories, an accent classifier is used to identify the accents of the speech audio in the speech library and automatically mark non-target accent speech as non-standard speech.

[0046] For a speech library that stores only a single accent, speech audio is randomly extracted directly from the speech library corresponding to the non-target accent as non-standard speech.

[0047] In one embodiment, the audio data requirements for non-standard and standard speech can be freely set by the user according to actual needs. For example, a single speech segment may be required to be 5-10 seconds long, with a sampling rate of 16kHz, and include speakers of different genders, ages, and speaking speeds to ensure accent diversity.

[0048] In one embodiment, standard speech can be extracted from a standard speech database. Specifically, publicly available target accent databases, such as American English speech databases or multi-speaker British English speech databases, are used directly to extract clear, accent-free standard speech segments.

[0049] In another embodiment, standard speech can also be obtained through TTS synthesis. Specifically, random text is sampled from publicly available text corpora (such as Wikipedia or news texts) to obtain random text unrelated to non-standard speech content. Standard speech is then synthesized using a TTS system with the target accent as input to the obtained random text.

[0050] Further, obtaining at least one standard speech corresponding to the target accent includes: extracting at least one standard speech from the standard speech database corresponding to the target accent; or, randomly sampling a public text corpus to obtain random text, and performing speech synthesis based on the target accent and the random text to obtain at least one standard speech corresponding to the target accent.

[0051] In one embodiment, standard speech can be extracted from a standard speech library. A high-quality single-speaker or multi-speaker standard speech library corresponding to the target accent is selected, and standard speech is extracted from it according to the audio data requirements.

[0052] The audio data requirements can be freely set by the user according to their actual needs. For example, the voice quality requirements include a signal-to-noise ratio of >30dB and no background noise; the voice content requirements include the speech of common words, phrases and sentences, covering different speech rates (such as 120-180 syllables / minute) and prosodic patterns (such as statements, questions and exclamations).

[0053] During the speech extraction process, a speech activity detection tool is used to extract valid speech segments and remove silence and non-speech parts (such as breathing sounds and equipment noise).

[0054] After extraction, the extracted speech audio is standardized. The standardization steps include: unifying the sampling rate; normalizing the volume to avoid volume differences affecting subsequent model training; and text alignment verification to check the text annotations corresponding to the speech to ensure that phonemes and pronunciations are strictly matched (e.g., excluding samples with incorrect annotations of elision or connected speech).

[0055] In another embodiment, standard speech can also be obtained through TTS synthesis. Specifically, random text is obtained by randomly sampling from a publicly available text corpus.

[0056] Public text corpora can include Wikipedia, news corpora, movie scripts, social media conversations, etc.

[0057] The random sampling mechanism can be stratified sampling, stratifying by text length (e.g., short sentences < 10 words, medium sentences 10-20 words, long sentences > 20 words), and randomly sampling a preset proportion of samples from each stratum to ensure coverage of different sentence structures. Then, deduplication and cleaning are performed, using the SimHash algorithm to remove duplicate text and filtering out sentences containing special symbols, foreign words, or grammatical errors to obtain the final random text.

[0058] In one embodiment, random text is input into a TTS (Text-to-Speech) system, and the TTS generates standard speech corresponding to the target accent based on preset synthesis parameters.

[0059] The synthesis parameters include speech style and speech rate (e.g., the intensity of the retroflexion of the "r" sound in American English, the glottal stop feature of the "t" sound in British English; the fundamental frequency range of American English is 100-500Hz, and the speech rate is 150 syllables / minute; the fundamental frequency range of British English is 90-450Hz, and the speech rate is 140 syllables / minute), the speaker's voice timbre, and the output format.

[0060] In another embodiment, the standard speech generated by the TTS model undergoes quality screening. Low-quality samples are filtered out using objective quality screening metrics and / or subjective human review (annotation team of at least three people assessing accent purity). The quality screening metrics can be freely set by the user according to actual needs.

[0061] In the above embodiments, standard speech is synthesized through a TTS model, eliminating the need for manual recording. This allows for the rapid generation of large-scale standard speech, constructing a large-scale and diverse target accent standard speech library, solving the problem of scarce parallel corpora, and reducing the cost of obtaining standard speech.

[0062] S202. Discretize each of the non-standard speech and each of the standard speech to obtain a discretized first speech unit sequence and a second speech unit sequence. In one embodiment, a pre-trained Hubert model is used to extract intermediate hidden representations of non-standard and standard speech, such as the output of the 6th layer of the Hubert model. The Hubert model learns phoneme-level features through a Masked Prediction task, and its output Hubert features contain phoneme-level information but do not include non-core features such as accents and prosody.

[0063] K-means clustering was performed on the HuBERT features of both non-standard and standard speech, mapping continuous features to a finite set of discrete tokens (e.g., 1000 cluster centers). Each speech frame was assigned a clustering index, forming a sequence of discrete speech units. Simultaneously, an accent encoder was used to extract accents from the non-standard speech, obtaining non-standard accent features.

[0064] In one embodiment, the first speech unit sequence is a token sequence obtained by discretizing non-standard speech. The second speech unit sequence is a token sequence obtained by discretizing standard speech.

[0065] S203. The pre-trained model is trained based on the first speech unit sequence and the second speech unit sequence to obtain the accent conversion model.

[0066] In one embodiment, a Transformer encoder-decoder architecture is used as the basic model architecture. A first sequence of non-standard speech units is randomly paired with a second sequence of standard speech units to form non-parallel training samples, without requiring content alignment.

[0067] The loss function includes the cross-entropy loss function, which is used to minimize the difference between the transformed token sequence and the second speech unit sequence corresponding to the standard speech.

[0068] In a specific embodiment, the encoder-decoder basic model architecture based on Transformer takes a non-standard token sequence and non-standard accent features as input and outputs a mapped standard token sequence (the token corresponding to the target accent).

[0069] The cross-entropy loss is used to calculate the loss value between the standard token sequence and the second speech unit sequence output by the model, in order to optimize the accent conversion model based on the Transformer encoder-decoder, so that the distribution of its output native token sequence is close to that of the standard token sequence.

[0070] In the above embodiments, the accent conversion model is trained with non-parallel data, which can quickly generate large-scale standard speech, solve the problem of scarce parallel corpora, and reduce the cost of obtaining standard speech.

[0071] Please see Figure 4 , Figure 4This is a schematic flowchart illustrating an accent conversion method based on discrete speech representation provided in an embodiment of this application. This accent conversion method based on discrete speech representation can be applied in a server to obtain a speech unit sequence containing duration information through a stream matching duration prediction model. It then iteratively obtains a Mel spectrum by combining the speech unit sequence containing duration information with timbre features, and further encodes the Mel spectrum to obtain the target speech. This achieves the conversion from discrete token sequences to natural speech waveforms, improving the pronunciation accuracy of the target speech and the natural fluency of the speech rhythm, while preserving the unique timbre of the original speaker.

[0072] like Figure 4 As shown, step S103 of the accent conversion method based on discrete speech representation specifically includes steps S301 to S303.

[0073] S301. Obtain the total duration of the speech to be converted, and based on the stream matching duration prediction model, predict the duration of each speech unit in the target speech unit sequence according to the total duration to obtain a speech unit sequence containing duration information. In one embodiment, in some scenarios, it is necessary to strictly align the time sequence of the converted speech with that of the original speech audio, such as in movie dubbing. Because the audio needs to match the performer's lip movements, the time sequence of the converted speech audio needs to be strictly aligned with the time of the original audio. Therefore, this embodiment introduces a total duration constraint to ensure that the time information of the generated target speech is strictly aligned with that of the speech to be converted.

[0074] In one embodiment, the original total duration of the speech to be converted is obtained as a constraint for duration prediction.

[0075] like Figure 2 As shown, the non-autoregressive synthesizer includes a stream-matching duration prediction model and a stream-matching acoustic model. The non-autoregressive synthesizer receives a speaker embedding vector representing timbre features and a total duration, and uses the target speech unit sequence and the total duration as input parameters to the stream-matching duration prediction model.

[0076] The stream matching duration prediction model encodes the target speech unit sequence using a Transformer encoder, captures the contextual dependencies between tokens, and learns a mapping function from an initial random duration distribution to the target duration distribution through a decoder. The noise is iteratively adjusted to a duration parameter sequence that conforms to the total duration through a mapping function, and then the duration parameter sequence is aligned with the target speech unit sequence to obtain a speech unit sequence containing duration information. Let t represent the duration state at time t, where t∈[0,1] is the time step.

[0077] In one embodiment, compared to traditional rule-based duration allocation (such as fixed phoneme duration templates), the stream matching model learns the duration distribution patterns of real speech to generate rhythms that better conform to human pronunciation habits. It supports switching between two modes: strict total duration alignment (such as matching lip movements in film and television dubbing) and natural rhythm priority (such as dialogues in financial and medical voice assistants).

[0078] In another embodiment, duration control can also adjust the speech rate of the target speech after it has been obtained. Specifically, a scaling factor is calculated based on the total speech duration of the speech to be converted and the target speech. If the scaling factor value is greater than 1, it indicates that the synthesized target speech is too fast and needs to be slowed down; if the scaling factor value is less than 1, it indicates that the synthesized target speech is too slow and needs to be slowed up. Here, scaling factor = speech duration of speech to be converted / speech duration of target speech.

[0079] S302. Based on the stream matching acoustic model, the Mel spectrum is obtained by iterating according to the speech unit sequence containing duration information and the timbre features. In one embodiment, a sequence of speech units containing duration information and a speaker embedding vector representing timbre features are input into a stream-matching acoustic model. The stream-matching acoustic model uses the frame-level speech unit sequence as the content condition and the timbre features as the timbre condition to iteratively generate a high-quality Mel spectrum starting from noise.

[0080] In a specific embodiment, the speech unit sequence containing duration information and timbre features are fused through a cross-attention mechanism to generate a conditional vector z, which contains phonemes, duration, and timbre features.

[0081] The flow matching generation model uses the random noise Mel spectrum as the initial state and predicts the velocity field based on the condition vector z. ,pass The noise Mel spectrum is gradually optimized to evolve into the spectral characteristics of the target accent and the original timbre, thus obtaining the final Mel spectrum.

[0082] S303. The Mel spectrum is encoded using a speech encoder to obtain the target speech.

[0083] In one embodiment, the speech encoder may employ a HiFi-GAN vocoder, including a generator that upsamples the Mel spectrum to the waveform sampling rate through stacked transposed convolutional layers and recovers high-frequency details through residual blocks; and a multi-scale discriminator that distinguishes the generated waveform from the real speech from different time resolutions (e.g., 20ms, 50ms windows) to drive the generator to optimize sound quality.

[0084] Specifically, the generator converts the Mel spectrum into a time-domain waveform signal and enhances the naturalness of the waveform through a non-linear activation function. Then, sound quality is optimized through adversarial training and feature matching loss. Specifically, adversarial training minimizes the adversarial loss between the generator and the discriminator, improving the realism of the waveform. Feature matching loss matches the generated waveform with features such as the Mel spectrum and MFCC (Mel-Frequency Cepstral Coefficients) of real speech, ensuring that pronunciation details (such as stress positions and intonation inflections) conform to the target accent.

[0085] In the above embodiments, a speech unit sequence containing duration information is obtained through a stream matching duration prediction model. A Mel spectrum is obtained by iteratively combining the speech unit sequence containing duration information and timbre features through a stream matching generation model. Then, the Mel spectrum is encoded into speech by a speech encoder to obtain the target speech. This realizes the conversion from discrete token sequence to natural speech waveform, improves the pronunciation accuracy of the target speech and the natural fluency of the speech rhythm, and preserves the unique timbre of the original speaker.

[0086] Further, step S103 also includes: acquiring the target style data of the accent to be converted; adjusting each speech unit in the target speech unit sequence according to the target style data based on the stream matching style adaptation model to obtain a speech unit sequence containing style parameters; iterating according to the speech unit sequence containing style parameters and the timbre features based on the stream matching generation model to obtain the Mel spectrum; and performing speech encoding on the Mel spectrum based on the speech encoder to obtain the target speech.

[0087] In one embodiment, target style data refers to a set of acoustic parameters that control the style of speech expression. For example, standard American English can simultaneously support styles such as formal broadcasting, daily conversation, and emotional reading.

[0088] It provides a style control interface that allows users to adjust the style intensity or specify the target style through style parameter sliders or text descriptions (such as broadcasting like a news anchor) and maps the text descriptions to quantization parameters.

[0089] In one embodiment, the target speech unit sequence is fused with the target style parameters to generate a structured sequence containing style information, providing style constraints for Mel spectrum generation.

[0090] Specifically, the target style data is transformed into a low-dimensional style embedding vector, such as by processing parameters like the F0 curve and speech rate through a Transformer encoder, and outputting a 256-dimensional style vector.

[0091] The stream matching style adaptation model concatenates the style embedding vector with the token features of the target speech unit sequence, and achieves the fusion of style and phoneme information through a conditional fusion layer.

[0092] In a flow-matching style adaptation model, the decoder learns a mapping function from random noise to the distribution of target style parameters. , where h is the feature after fusing style and target speech unit sequences, and t∈[0,1] is the iteration time step.

[0093] The noise is iteratively adjusted to conform to the target style through a mapping function, and the style parameter sequence is aligned with the target speech unit sequence to form a speech unit sequence containing style parameters.

[0094] In one embodiment, a sequence of speech units containing style parameters and timbre features are fused using a cross-attention mechanism to generate a conditional feature vector. It also includes phonemes, style, and timbre characteristics.

[0095] The flow matching generation model randomly samples noise Mel-frequency spectra that follow a Gaussian distribution. The generator predicts the velocity field based on conditional eigenvectors. The spectrum is gradually optimized using the following iterative formula: During the iteration process, a style loss function is used to ensure that the style parameters of the generated spectrum are consistent with the target.

[0096] In one embodiment, the speech encoder converts the generated Mel spectrum into a time-domain waveform through a transposed convolutional layer, while optimizing the sound quality through a multi-scale discriminator, and finally outputs the target speech that integrates the target accent, target style, and original timbre.

[0097] In the above embodiments, while converting accents, the target style (such as emotion, speech rate, and rhythm) is flexibly adapted, and the original speaker's timbre is preserved, which improves the flexibility of accent conversion and thus improves the adaptability of the target speech to the target scene.

[0098] Please see Figure 5 , Figure 5 This application provides a schematic block diagram of an accent conversion device based on discrete speech representation, which is used to perform the aforementioned accent conversion method based on discrete speech representation. The accent conversion device can be configured on a server.

[0099] like Figure 5 As shown, the accent conversion device 400 based on discrete speech representation includes: The speech discretization processing module 401 is used to discretize and extract accents from the speech to be converted when it receives the speech to be converted, so as to obtain a discretized initial speech unit sequence and original accent features. The target accent conversion module 402 is used to perform autoregressive conversion on the initial speech unit sequence and the original accent features based on a preset accent conversion model to obtain the target speech unit sequence corresponding to the target accent. The target speech acquisition module 403 is used to acquire the timbre features corresponding to the speech to be converted, and synthesize the target speech corresponding to the target accent based on the timbre features and the target speech unit sequence.

[0100] Furthermore, the accent conversion device based on discrete speech representation further includes an accent conversion model acquisition module, which includes: The speech acquisition unit is used to acquire at least one non-standard speech corresponding to the non-target accent and at least one standard speech corresponding to the target accent; A discrete processing unit is used to discretize each of the non-standard speech and each of the standard speech to obtain a discretized first speech unit sequence and a second speech unit sequence. The model training unit is used to train the pre-trained model based on the first speech unit sequence and the second speech unit sequence to obtain the accent conversion model.

[0101] Furthermore, the speech acquisition unit is specifically used to extract at least one standard speech from the standard speech database corresponding to the target accent; or, to randomly sample publicly available text corpus to obtain random text, and to perform speech synthesis based on the target accent and the random text to obtain at least one standard speech corresponding to the target accent.

[0102] Furthermore, the target speech acquisition module 403 includes: The duration prediction unit is used to obtain the total duration of the speech to be converted, and based on the stream matching duration prediction model, to predict the duration of each speech unit in the target speech unit sequence according to the total duration, so as to obtain a speech unit sequence containing duration information. The Mel spectrum acquisition unit is used to obtain the Mel spectrum by iterating according to the speech unit sequence containing duration information and the timbre features based on the stream matching acoustic model. The target speech acquisition unit is used to perform speech encoding on the Mel spectrum based on the speech encoder to obtain the target speech.

[0103] Furthermore, the target speech acquisition module 403 further includes: The style adjustment unit is used to acquire the target style data of the accent to be converted, and based on the stream matching style adaptation model, adjust each speech unit in the target speech unit sequence according to the target style data to obtain a speech unit sequence containing style parameters. The Mel spectrum acquisition unit is used to obtain the Mel spectrum by iterating according to the speech unit sequence containing style parameters and the timbre features based on the stream matching generation model, and to perform speech encoding on the Mel spectrum based on the speech encoder to obtain the target speech.

[0104] Furthermore, the speech discrete processing module 401 includes: The discrete processing unit is used to discretize the speech to be converted based on a preset self-supervised learning model when the speech to be converted is received, to obtain discrete speech units, and to cluster the discrete speech units to obtain the initial speech unit sequence. The accent extraction unit is used to extract accents from the speech to be converted based on a preset accent editor, so as to obtain the original accent features.

[0105] Furthermore, the accent conversion device 400 based on discrete speech representation further includes a speech-to-be-converted determination module, which includes: An accent recognition unit is used to, upon receiving a voice audio, perform accent recognition on the voice audio based on an accent classification model, and determine the accent category corresponding to the voice audio. The speech to be converted determination unit is used to determine the audio with an accent category that is not the target accent as the speech to be converted.

[0106] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0107] The aforementioned apparatus can be implemented as a computer program, which can be used in, for example... Figure 6 It runs on the computer device shown.

[0108] Please see Figure 6 , Figure 6 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.

[0109] See Figure 6 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0110] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any accent conversion method based on discrete speech representation.

[0111] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0112] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When executed by a processor, the computer program enables the processor to perform any accent conversion method based on discrete speech representation.

[0113] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0114] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0115] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: Upon receiving the speech to be converted, the speech is discretized and accent extracted to obtain a discretized initial speech unit sequence and original accent features. Based on a preset accent conversion model, an autoregressive conversion is performed on the initial speech unit sequence and the original accent features to obtain the target speech unit sequence corresponding to the target accent. The timbre features corresponding to the speech to be converted are obtained, and the target speech corresponding to the target accent is synthesized based on the timbre features and the target speech unit sequence.

[0116] In one embodiment, before the processor performs an autoregressive transformation on the initial speech unit sequence and the original accent features based on a preset accent conversion model to obtain the target speech unit sequence corresponding to the target accent, it is further configured to: Obtain at least one non-standard speech corresponding to the non-target accent and at least one standard speech corresponding to the target accent; Discretize each of the non-standard speech and each of the standard speech to obtain a discretized first speech unit sequence and a second speech unit sequence; The accent conversion model is obtained by training the pre-trained model based on the first speech unit sequence and the second speech unit sequence.

[0117] In one embodiment, when the processor acquires at least one standard speech corresponding to the target accent, it is configured to: Extract at least one standard speech record from the standard speech database corresponding to the target accent; or, Randomly sample publicly available text corpora to obtain random text, and perform speech synthesis based on the target accent and the random text to obtain at least one standard speech corresponding to the target accent.

[0118] In one embodiment, when the processor synthesizes target speech corresponding to the target accent based on the timbre features and the target speech unit sequence, it is configured to: The total duration of the speech to be converted is obtained. Based on the stream matching duration prediction model, the duration of each speech unit in the target speech unit sequence is predicted according to the total duration to obtain a speech unit sequence containing duration information. Based on the stream matching acoustic model, the Mel spectrum is obtained by iterating according to the speech unit sequence containing duration information and the timbre features. The target speech is obtained by encoding the Mel spectrum using a speech encoder.

[0119] In one embodiment, before performing speech encoding on the Mel spectrum based on the speech encoder to obtain the target speech, the processor is further configured to perform: Obtain the target style data of the accent to be converted, and based on the stream matching style adaptation model, adjust each speech unit in the target speech unit sequence according to the target style data to obtain a speech unit sequence containing style parameters; Based on the stream matching generation model, the Mel spectrum is obtained by iterating according to the speech unit sequence containing style parameters and the timbre features, and the Mel spectrum is then encoded into speech based on the speech encoder to obtain the target speech.

[0120] In one embodiment, when the processor receives speech to be converted, it performs discretization processing and accent extraction on the speech to be converted to obtain a discretized initial speech unit sequence and original accent features, and implements the following: Upon receiving the speech to be converted, the speech is discretized based on a preset self-supervised learning model to obtain discrete speech units, and the discrete speech units are clustered to obtain the initial speech unit sequence. Based on a preset accent editor, the accent of the speech to be converted is extracted to obtain the original accent features.

[0121] In one embodiment, before the processor discretizes and extracts accents from the received speech to be converted to obtain a discretized initial speech unit sequence and original accent features, it is further configured to: Upon receiving a voice audio file, the voice audio file is subjected to accent recognition based on an accent classification model to determine the accent category corresponding to the voice audio file. Audio with an accent category other than the target accent is identified as the speech to be converted.

[0122] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the accent conversion methods based on discrete speech representation provided in the embodiments of this application.

[0123] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0124] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An accent conversion method based on discrete speech representation, characterized in that, include: Upon receiving the speech to be converted, the speech is discretized and accent extracted to obtain a discretized initial speech unit sequence and original accent features. Based on a preset accent conversion model, an autoregressive conversion is performed on the initial speech unit sequence and the original accent features to obtain the target speech unit sequence corresponding to the target accent. The timbre features corresponding to the speech to be converted are obtained, and the target speech corresponding to the target accent is synthesized based on the timbre features and the target speech unit sequence.

2. The accent conversion method based on discrete speech representation according to claim 1, characterized in that, Before obtaining the target speech unit sequence corresponding to the target accent by performing autoregressive transformation on the initial speech unit sequence and the original accent features based on the preset accent conversion model, the method further includes: Obtain at least one non-standard speech corresponding to the non-target accent and at least one standard speech corresponding to the target accent; Discretize each of the non-standard speech and each of the standard speech to obtain a discretized first speech unit sequence and a second speech unit sequence; The accent conversion model is obtained by training the pre-trained model based on the first speech unit sequence and the second speech unit sequence.

3. The accent conversion method based on discrete speech representation according to claim 2, characterized in that, The acquisition of at least one standard speech record corresponding to the target accent includes: Extract at least one standard speech record from the standard speech database corresponding to the target accent; or, Randomly sample publicly available text corpora to obtain random text, and perform speech synthesis based on the target accent and the random text to obtain at least one standard speech corresponding to the target accent.

4. The accent conversion method based on discrete speech representation according to claim 1, characterized in that, The process of synthesizing the target speech corresponding to the target accent based on the timbre features and the target speech unit sequence includes: The total duration of the speech to be converted is obtained. Based on the stream matching duration prediction model, the duration of each speech unit in the target speech unit sequence is predicted according to the total duration to obtain a speech unit sequence containing duration information. Based on the stream matching acoustic model, the Mel spectrum is obtained by iterating according to the speech unit sequence containing duration information and the timbre features. The target speech is obtained by encoding the Mel spectrum using a speech encoder.

5. The accent conversion method based on discrete speech representation according to claim 1, characterized in that, The process of synthesizing the target speech corresponding to the target accent based on the timbre features and the target speech unit sequence further includes: Obtain the target style data of the accent to be converted, and based on the stream matching style adaptation model, adjust each speech unit in the target speech unit sequence according to the target style data to obtain a speech unit sequence containing style parameters; Based on the stream matching generation model, the Mel spectrum is obtained by iterating according to the speech unit sequence containing style parameters and the timbre features, and the Mel spectrum is then encoded into speech based on the speech encoder to obtain the target speech.

6. The accent conversion method based on discrete speech representation according to claim 1, characterized in that, Upon receiving the speech to be converted, the process involves discretizing and extracting the accent from the speech to obtain a discretized initial speech unit sequence and original accent features, including: Upon receiving the speech to be converted, the speech is discretized based on a preset self-supervised learning model to obtain discrete speech units, and the discrete speech units are clustered to obtain the initial speech unit sequence. Based on a preset accent editor, the accent of the speech to be converted is extracted to obtain the original accent features.

7. The accent conversion method based on discrete speech representation according to any one of claims 1 to 6, characterized in that, Before performing discretization and accent extraction on the received speech to be converted to obtain a discretized initial speech unit sequence and original accent features, the process further includes: Upon receiving a voice audio file, the voice audio file is subjected to accent recognition based on an accent classification model to determine the accent category corresponding to the voice audio file. Audio with an accent category other than the target accent is identified as the speech to be converted.

8. An accent conversion device based on discrete speech representation, characterized in that, include: The speech discretization module is used to discretize and extract accents from the received speech to be converted, thereby obtaining a discretized initial speech unit sequence and original accent features. The target accent conversion module is used to perform autoregressive conversion on the initial speech unit sequence and the original accent features based on a preset accent conversion model to obtain the target speech unit sequence corresponding to the target accent. The target speech acquisition module is used to acquire the timbre features corresponding to the speech to be converted, and synthesize the target speech corresponding to the target accent based on the timbre features and the target speech unit sequence.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, in executing the computer program, implement the accent conversion method based on discrete speech representation as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to implement the accent conversion method based on discrete speech representation as described in any one of claims 1 to 7.